跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

When Two Top Companies Make Opposite Bets, Where Does AI Memory End Up?

Original · Unique Research · 2026-08-04

Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening essay, four insight sections, closing reflection, and the full roundtable transcript. All named companies, products, and people are preserved. Cost figures, star counts, funding amounts, and product capabilities are speaker claims, not independently verified findings.

AI Industry Observer

"Text Memory Will Be Eaten by the Model": Two Memory-Infrastructure Companies Bet on Opposite Futures

"At the end of the day, what's being competed for is judgment — what to remember, what to forget, whether to believe."

Memory is one of the hottest words in AI this half-year, but the industry has yet to reach consensus on what memory actually is.

At a roundtable on agent memory at the Unique Awards, four guests — EverMind incubated by Shanda, Silicon Valley startup Memories.ai, cybersecurity firm Wanjing Security, and education technology company Shiji Tianhong (世纪天鸿) — took the discussion from underlying technology to vertical scenarios and covered it thoroughly. The most interesting part is not that they reached consensus, but that the two memory-infrastructure companies made bets in completely opposite directions. Behind this lies a question everyone building "AI infrastructure" should think through: will the layer you're building be eaten by the model itself?

A Bold Call: Text Memory Will Ultimately Be Eaten by the Model

Memories.ai founder Shen Junxiao (沈俊潇), a former Meta research scientist, settled on one conclusion two years ago: text memory will ultimately be subsumed by the model itself.

His logic goes like this: when ChatGPT first appeared three years ago, the industry's understanding of memory was simply "add a layer of RAG." Then MemGPT (later Letta) began exploring personal memory. More recently, OpenClaw and Claude Code have moved to a pure Markdown index file system — essentially giving the model a library catalog, telling it where everything is, and letting it find things itself. The driving force behind this evolutionary path is not progress in memory technology itself, but that models have longer and longer context windows and stronger and stronger capabilities. In other words, the memory layer you build on top of the model (the "agent harness") will grow thinner as models get stronger.

"If this judgment holds, then for a company that wants to build 'infrastructure' rather than 'applications,' betting on text memory is dangerous — this layer will be continuously eroded by the model itself, and the ones who do it well are usually not infrastructure companies but application providers who truly understand the business."

That is why Memories.ai never did text memory from day one — they went all in on visual memory. The reasoning is direct: video is not a native input to large models; its information entropy is extremely high, so it must first be compressed and encoded into structured information before models can understand and invoke it. This layer is sufficiently "infrastructure" and does not grow thinner with business changes. Over two years, they compressed video encoding from ten embeddings to one embedding, reducing encoding costs by at least 100x and storage costs by at least 50x. Today, storing 1,000 hours of video costs only 20 RMB, and encoding costs only $100. They also have strategic partnerships with Qualcomm and NVIDIA, deploying encoding models directly on local cameras or terminals — the video itself never goes to the cloud; only desensitized structured information enters the cloud "video data lake." Shen positions this as the first step toward physical AI: robots are not the starting point for physical AI, but letting AI first understand the real world through cameras is.

Another Bet: Simply Nail "General-Purpose Memory" Thoroughly

EverMind took the completely opposite path — they bet that "general-purpose memory" can actually be built, and that the text/structured-data route is far from exhausted.

The company behind it is Shanda Group. COO Han Yunyun's (韩云芸) logic carries a philosophical color: they decompose intelligence into two modules — reasoning and memory. Reasoning corresponds to "logical intelligence"; memory corresponds to "subjective intelligence." Shanda published a white paper on long-term memory two years ago, predicting that agents would become AI's best application scenario — a prediction that has been validated today.

EverMind's approach is to build a four-layer architecture bottom-up: the bottom layer is the interface layer, responsible for ingesting heterogeneous data; above that is the memory layer, which structures raw data and supports CRUD operations; above that is the index layer, specifically addressing retrieval accuracy, latency, and cost — they even trained dedicated models for this; the top layer is the agent layer, connecting to applications and deliberately leaving custom extension space, relying on the open-source community to feed back use cases. This system has already incubated several products: the open-source framework EverOS surpassed 10,000 GitHub stars; Raven, which packages memory with 100,000 skills into a complete harness, surpassed 1,000 stars within three days of launch and already achieves code-level rewriting, continuously optimizing invocation paths as usage grows; and EverMe, built for ordinary users without programming ability. Their latest move is integrating "reasoning" into the memory system itself, achieving continuous memory updates across 400-step reasoning processes — meaning agents won't "amnesiate" mid-task on long-chain work, eliminating the need to start over.

"Shen Junxiao and Han Yunyun know each other privately, and each praises the other as 'having chosen a path of great value.' But the divergence between the two paths shows that the memory sector is far from convergence. Whether you bet on the model's ceiling or the model's floor determines completely different products and business models."

The Real Pain Point Isn't Whether the Model Is Smart Enough — It's Whether You Dare Trust It

If the two infrastructure companies were debating "what to build," the two vertical-scenario application guests revealed why these elegant memory architectures fall short in real business.

Wanjing Security founder Si Hongxing (司红星) described a stark reality: their clients are national critical infrastructure entities in power, energy, and finance. Data must be deployed offline; they cannot use the strongest models like Claude or GPT-5.6 — only 27B-class open-source small models. When model capability is insufficient, long-chain tasks easily fall into "spin" — repeatedly calling one tool without breaking out, even starting to hallucinate. Their solution is a reflection and spin-detection mechanism: once abnormal behavior is detected, force interruption, search the memory bank for "whether similar problems have been encountered historically," reinject past experience into context, and continue. Memory itself has a scoring mechanism, scoring each entry across eight dimensions — relevance, timeliness, user preference, and others — with older memories weighted lower and recalled with weighted scoring.

"He admitted frankly: agents are not yet truly deployed on the front lines of cybersecurity; most remain in the laboratory stage — the reason is not that technology is insufficient, but that these clients 'do not allow trial and error.' Once an agent makes the wrong call and closes an external network interface, it could directly affect residents' electricity and water supply. In industries with near-zero fault tolerance, the core value of a memory system is not 'smarter' but 'more controllable.'"

Shiji Tianhong's Zhang Minsong (张民松) gave a nearly isomorphic answer from the education scenario. He summarized AI application bottlenecks in teaching into two words: boundary and credibility. Boundary refers to the sequential and staged nature of teaching — the same problem, a third-grader can only solve it using arithmetic methods, not prematurely using linear equations, even though the latter is "smarter" — in a teaching context, that is overreach and the wrong answer. Credibility refers to content compliance and subject accuracy; compliance is also an educational ethics requirement. On subject accuracy, models generally show a distribution pattern: higher grade levels perform worse than lower grade levels, and humanities perform better than STEM. For example, mathematics is basically usable at the elementary level, prone to errors in middle school, and by high school it is basically "beyond recognition." Their current solution is to inject content boundaries as constraints directly into the memory layer, and rely on their own 30 years of accumulated foundational content to reduce model output variance, rather than depending solely on the model's own capability.

At the End of the Day, What's Being Competed For Is the "Judgment" Layer

Putting the four guests' judgments together, this discussion of memory appears on the surface to argue about "how memory should be stored, whether to use RAG or Markdown," but the real divergence and value both land on the same thing: what to remember, what to forget, whether to believe — this judgment itself is what is scarce.

Shen Junxiao judges that "this layer of judgment should be handed to business parties; infrastructure should move toward areas like vision that models can't yet reach"; Han Yunyun judges that "a general-purpose memory architecture itself is worth building, and can in turn enable reasoning and memory to co-evolve"; Si Hongxing's judgment manifests in "when to forcibly interrupt an agent, where to find similar experience"; Zhang Minsong's judgment manifests in "can this problem be taught now, can this statement be made." The reasoning capability of models is growing stronger and cheaper — this is nearly an industry consensus — but who tells the model "be humble on this matter, be stubborn on that matter" is the position this sector is actually competing for.

More Conversation Details

Speakers

EverMind COO Han Yunyun (韩云芸)

Wanjing Security Founder & Chairman Si Hongxing (司红星)

Memories.ai CEO Shen Junxiao (沈俊潇)

Shandong University & Shiji Tianhong Distinguished Researcher & AI Lead Zhang Minsong (张民松)

Host

Unique Capital Partner Zhao Liang Abner (赵亮)

Zhao Liang:

Today's topic is how agents form continuous context — the topic of agent memory, which has been extremely hot recently. We have four guests from different fields. First, please each introduce yourself in one to two minutes, covering your company's product, service, and target customers.

Han Yunyun:

Hello everyone, I'm from Shanda Group. Shanda went all in on AI three years ago and has been incubating its own AI businesses. Currently, my partner Ya Feng and I mainly lead the team working on AI memory. We've also recently started combining memory and reasoning technologies. From a brain science perspective, Shanda decomposes human intelligence into two major modules: reasoning and memory. The two teams have been working on merging their tracks. Our products are mainly EverOS, Raven, and EverMe. Please check our website and GitHub. EverOS recently passed 10,000 stars — thank you all for your support.

Si Hongxing:

Hello everyone, I'm Si Hongxing from Wanjing Security. Our company mainly does cybersecurity, serving large B2B clients — extremely large B2B. Clients are primarily national critical infrastructure entities with rigid needs in cybersecurity, such as power, energy, finance, and telecom operators. Our vision and mission is to use AI to complete the intelligent transformation of cyberattack and defense. Our AI products include MemorFit, and on top of the AI foundation, we've developed a new programming language — a DSL built for cybersecurity. Combining this development language with an AI brain enables intelligent attack and defense. At the memory layer, we discovered during our work, or rather during agent construction, that memory is extremely important. I'm very happy to be here to discuss memory with everyone.

Shen Junxiao:

Hello everyone, my name is Shen Junxiao, founder and CEO of Memories.ai. I was previously a research scientist at Meta. Most of our team comes from Meta and the Meta Superintelligence Lab. We're a Silicon Valley startup serving companies that have many cameras or large amounts of video data. We do visual memory, which is different from text memory — text memory is built on large amounts of textual data, while ours is built on large amounts of visual and video data. What we do is turn video into structured information that AI can understand and invoke. You can think of us as the video version of Snowflake — converting video into a video data lake, and then doing three things on that video database: first, get insights; second, build agent workflows; third, do data curation to train models. That's our visual memory.

Zhao Liang:

Memories was just founded?

Shen Junxiao:

That's right, founded a year and a half ago, raised over ten million dollars in Silicon Valley.

Zhao Liang:

Very young entrepreneurs. You currently have strategic partnerships with Qualcomm, NVIDIA, and Microsoft. During Microsoft's CEO talk at Microsoft Build, your logo was right behind him.

Shen Junxiao:

Correct.

Zhang Minsong:

Hello everyone, I'm Zhang Minsong, AI lead at Shiji Tianhong, and I also do applied exploration of mathematical models at Shandong University. Shiji Tianhong is an established listed educational publishing company with over 30 years of history, having accumulated massive amounts of content. We began exploring AI + education scenarios in 2023, and have since launched AI products like Xiao Hong Teaching Assistant, primarily empowering schools across teaching, learning, research, and assessment. I'm glad to be here to share and exchange ideas.

Zhao Liang:

Mr. Zhang is also from an internet background, having worked in technology before transitioning to the education sector and deeply developing AI applications in education scenarios. Thank you all for the introductions. Today, EverMind and Memories do underlying technology innovation, while Mr. Si and Mr. Zhang do scenario-based AI applications, all deeply focused on the topic of agent memory. So let's discuss memory from different angles.

How Should Memory Be Defined?

Zhao Liang:

My first question is for Mr. Shen, given the name "Memories." Definitions of memory in the market still vary significantly. Some people think ultra-long context embedded in the large model itself is memory. Others believe an external independent architecture, a separate memory layer, is needed. Others believe a hybrid architecture is necessary. From your perspective, how should memory be defined? Does it need to be layered?

Shen Junxiao:

When we were at Meta, we were actually working on the multimodal memory behind the personal assistant for Meta glasses. Our definition of memory back then was very different from today. That was over three years ago, right after ChatGPT came out. Then MemGPT came along — that company later became Letta, also a Silicon Valley startup. Their initial definition of AI memory was simply AI plus a layer of RAG — that was memory at the time. Then it gradually evolved into how to do personal memory, and then to OpenClaw, where personal memory shifted from RAG to a pure Markdown index file system. Before that, all those memories were for personalization, including Memories itself. But later, with the emergence of self-improving agents, memory began to be used for self-improvement.

So I think there are currently two uses of memory: the first is personalization, and the second is self-improvement. Of course, the memory I'm talking about is mainly the agent's own memory. There's also enterprise memory — a particularly representative company is Glean, a US startup that's already a unicorn. They convert all enterprise data into a vector database, turn it into an enterprise knowledge base, and make it available for agents to invoke. So memory is a very broad concept, encompassing enterprise agent memory, personal agent memory, the agent's own memory, and memory for self-improvement. And then there's the visual memory we do — visual memory for future physical AI. So memory itself is a big concept with various commercial paths and business models.

Zhao Liang:

Compared to three years ago, there have been tremendous changes and evolution.

Shen Junxiao:

Yes, these changes actually come from models getting stronger. In the past, you needed many RAG approaches to make memory better. Now, as model context grows and agent capabilities strengthen, often you only need to organize the indexing — the Markdown file system — and give it a library, tell it where everything is, and it can find things itself. Although it's token-intensive, accuracy is increasingly high. That's why OpenClaw and Claude Code both use pure Markdown to organize memory. But as memory grows — for example, as enterprise memory grows — you may still need a hybrid architecture: Markdown file format plus RAG. File format is token-efficient and accurate but token-intensive; RAG is less token-efficient — it's a spectrum. The more agentic, the more token-intensive but more accurate, and possibly more time-consuming. RAG is very fast, pure retrieval-based, but may not be as general or accurate. Generally, people now use a hybrid of RAG and Markdown for agent personal memory.

Why Is Memory Important? EverMind's Original Intent and Positioning

Zhao Liang:

My second question is for Ms. Han — why is memory important? EverMind has a very interesting background as a Shanda-incubated company. Chen Tianqiao has recently been very focused on foundational science and technology innovation, including brain neuroscience and AI for science. What was the original intent in founding EverMind? How does EverMind position itself? Please share based on your understanding of why memory is important.

Han Yunyun:

When I first joined Shanda, I was general manager of the Data Division at the Tianqiao Brain Science Research Institute, initially responsible for brain science. Mr. Chen mainly invests in two areas: brain science and artificial intelligence. We do both with the same starting point: technology still serves humanity. The ancient Greek saying is "man is the measure of all things" — we focus on the principles and methods of this "measurement." Shanda does brain science to use technology to help humanity discover the mysteries of the human brain and human intelligence; it does AI hoping that AI will help humanity explore the boundaries of cognition and extend our cognition. We do discovery-oriented intelligence, not generative intelligence. Generative intelligence may replace some human work, but we believe AI's greater potential is to help humans do what we currently cannot.

Returning to memory: we believe memory is another facet of human intelligence. While reasoning performs a kind of logical intelligence, memory performs subjective intelligence. In the future world, after AGI arrives, it should be a world of three-way interaction among humans, AI, and the world. So we want to use the best resources to build everything that helps humans prepare before AGI arrives, driving the evolution of intelligence from the memory dimension. About two years ago, Shanda published a white paper on long-term memory, predicting that two years later, agents might be AI's best interaction method or application scenario. So we started building all the infrastructure for agent memory at that time.

The first thing the EverMind team did was build a memory layer for internally incubated AI applications, including Tanka, the AI-native enterprise collaboration platform you may know. We help this complex scenario handle memory — user memory, agent memory, group memory, human collaboration, human-agent collaboration, and even agent-to-agent collaboration later. Through this approach, we find the paradigms and universality for providing memory to future agents, and then use technology at the underlying layer to help agents evolve further. This is our thread for doing memory. We have a different starting point from Memories.ai, but ultimately we arrive at the same destination. Two years later, memory has indeed become an industry hotspot. But logically speaking, the industry could foresee this: the deeper and more complex the use of LLMs, the more significant the value of memory becomes; when LLM applications move past turn-based chat into task execution, the value of memory surfaces.

We also believe memory and LLMs stand on the same vertical line. LLMs build a kind of general intelligence, omniscient intelligence, intelligence about the objective world. But memory builds a kind of subjective intelligence — that is, the "measurement" and convergence of the world from an individual or organizational perspective. A hot track in AI recently is world models. We also build models — ours can be analogized as a kind of "personality model," an intelligence that self-evolves based on having a certain subjectivity.

Cybersecurity Scenario: Why Is Memory Especially Important?

Zhao Liang:

Now let's turn to the two scenario companies. First, Mr. Si — the cybersecurity field has many complex decisions and ultra-long context. When you launched AI agents, you faced the same issues. In your scenario, why is memory especially important? What is the biggest pain point for your entire industry?

Si Hongxing:

One main characteristic of cybersecurity is long tasks — naturally long-chain tasks. Unlike using ChatGPT or Claude Code to do daily things like generating PPTs. For example, when an enterprise is attacked, there's only one lead. You have to run through all the tasks from various logs — various gateways, firewalls, logs coming in, possibly several gigabytes of files to analyze, many steps to call. In such long-chain tasks, you cannot do it without memory management.

Our memory manifests in three ways: first, short-term temporary memory injected into conversation context; second, mid-to-long-term task memory, because a single long-chain task involves many prompt inputs, and the task also needs to pull things back from long-term memory; third, long-term memory, which is an enterprise asset. Cybersecurity is generally team-based, not individual. Whether attacking or defending, how does an individual's experience get effectively stored? For example, during an attribution process, someone discovers a typical insight — which country the attacker organization belongs to, what its background is — that's a very typical memory. Once this person discovers it, how does the next person onboard natively reference it? Memory sharing is also a problem. All of this means cybersecurity agents naturally need structured memory management, not simply one task, one conversation, and done — that feels wasteful.

Zhao Liang:

Is this a pain point across the whole industry, or one Wanjing Security discovered?

Si Hongxing:

It's definitely needed across the cybersecurity industry — quite universal. Whether attacking or defending, whether cybersecurity or finance, any high-complexity long-chain task needs standardized memory management. But the cybersecurity industry is relatively behind. Compared to consumer-facing products, you're more advanced.

I say "behind" because our clients are national critical infrastructure entities — they cannot make mistakes, they do not allow trial and error, only small-scale trial and error is allowed, no large-scale direct deployment without validation. Because it involves memory security and agent security. If an agent makes a wrong decision and closes an external network interface, people might not even be able to pay their electricity bills — water and power could be cut off. That's serious. So agents are not yet widely deployed on the real front lines; they're all in the laboratory stage.

Differentiated Memory Requirements in Education

Zhao Liang:

Next, Mr. Zhang — Shiji Tianhong recently launched an education-track AI agent, Xiao Hong Teaching Assistant. At the memory layer, how do education scenarios differ in their memory requirements from other scenarios? Can the large model's own memory capability fully cover your current needs?

Zhang Minsong:

Technically, the implementation path is the same, just different application scenarios. The purpose of memory engineering is better decision-making. The education industry is special — it's a highly institutionalized, systematized, and organized scenario. Implementation is purposeful, planned, and organized behavior, which means the external decision space is relatively limited compared to other industries, and experience solidification is higher. For model applications, from a subject perspective we care about two things: first, boundary, meaning applicability; second, credibility. These two constrain the core bottleneck of AI deepening into education applications.

What is boundary? Education is a process of sequence and progression — there's order and stage. What to learn and when is fixed; you can't go ahead or fall behind. For example, the same problem — in third grade, solving for an unknown using two steps — you can't use linear equations in one variable; that's appropriate in fifth grade or middle school, otherwise it's overreach. To solve such problems, we inject boundary constraints into the underlying memory, whether dynamic or static — things like content application range, subject, grade level, curriculum standards — to limit the boundary.

What is credibility? We know that model reasoning is based on probability rather than actual logic, so hallucination rate is a matter of degree. When AI is applied to teaching, why should we trust it? It involves content compliance and accuracy. On compliance — for example, when asked about sensitive issues, such as the South China Sea Islands belonging to China, any hesitation or unnecessary explanation means non-compliance — that's unambiguous and non-negotiable. For subject accuracy, current general model output follows a distribution: higher grade levels perform worse than lower grade levels, and humanities outperform STEM. For example, mathematics is basically usable at the elementary level, very error-prone in middle school, and by high school it's basically unrecognizable. Our solution, whether memory or knowledge base retrieval, relies on solidified experience — i.e., content — to reduce model output variance. These are problems general models currently face in subject domains, and also problems and solution approaches we've encountered in our exploration and practice — for everyone's reference.

Zhao Liang:

It sounds like education scenarios don't simply pursue being smarter — they need to be more compliant, more accurate, more bounded. Memory is very important in this.

Zhang Minsong:

Correct.

Why Go All In on Visual Memory?

Zhao Liang:

Next question for Mr. Shen — Memories currently focuses on visual memory, but the entire AI ecosystem is built around text memory. What is the essential difference between visual memory and text memory? Why did you go all in on visual memory?

Shen Junxiao:

Two years ago when we started doing memory, we were certain about one thing: text memory will ultimately be covered by the model. The memory you build on it — text memory is, to some extent, an agent harness. As model capability grows stronger and stronger, the layer above will grow thinner and thinner. Since we want to be an infrastructure company, not an application company, we should build the core well. If the harness keeps getting thinner, it's hard to form a core, so we never wanted to do text memory from the start. Text is the native input of large language models. As model capability grows, the memory system you can build on top of it gets smaller and smaller, and those things are very business-oriented. Memory systems are important, but the ones who do text memory systems well will always be business parties, not infrastructure companies. Business parties truly understand the business — what the agent needs to do, how memory should be stored, how user memory should be managed — they know the business deeply, so they can do it well.

So we decided to do visual memory. It may be early, but vision is different. Human memory itself is visual. For example, when was the last time you went to the gym? Who did you see yesterday? The first thing that comes to mind isn't text — it's a fragment of the gym, and within that fragment you do reasoning. That aligns more with the future development path of visual models. And vision is not currently a native input of large models. Video is raw, with extremely high information entropy, containing all kinds of messy information. You need very foundational model technology to compress video, then do indexing or encoding, encoding it into structured information that AI can understand. This layer itself is already very infrastructure-like, and this encoding infrastructure doesn't change with the business.

Vision becomes video, becomes structured information, stored in a video data lake — that is visual memory. Once stored, what you do is make search and retrieval sufficiently accurate. Above that is the business layer. As an infrastructure company, what we do is how to turn video into structured data faster, better, and cheaper, and how to retrieve related structured data from tens of thousands of hours of video in a video database — like finding a needle in a haystack. Once found, how to organize it, form memory, and build the organizational layer and information agents want — that's the business layer. We work with large B2B clients through an FDE model, but our infrastructure is the layer below — video, video encoding, and the video data lake.

Zhao Liang:

Sounds extremely complex engineering-wise. What's the timeline for this?

Shen Junxiao:

We've been at it for two years, and so far we haven't seen many companies doing the same thing. We've done a lot of optimization on top. Video encoding sounds very heavy and expensive, but our video encoding model has taken the world model route from the start. We call it "only input" — it can take video, audio, text, context, and everything becomes embedding vectors. This model can now do one embedding — two years ago it was ten embeddings. One embedding has already reduced the entire video encoding cost by at least 100x, and video storage cost by at least 50x. These are optimizations we've done over two years.

Now when we talk to enterprise clients, they say video is expensive, storage is expensive, encoding is expensive. I say no, no — we've already done a lot of optimization. Now 1,000 hours of storage costs only 20 RMB, and 1,000 hours of encoding costs only $100. Very, very cheap. Qualcomm and NVIDIA are both strategic partners — we've been collaborating for two years. All encoding models can be placed locally. For example, a building can have a local compute center, or even all models can be placed directly on a phone. Video never goes to the cloud; after local encoding, only structured, desensitized information goes up to the cloud, forming a video data lake.

This is what we've always been certain about. I think this era may truly be transitioning to physical AI. The first step of physical AI arriving in the real world is actually not robots — it's letting AI first understand the real world through cameras, telling agents what's actually happening. For example, in retail, how do you know how many people came today, how many left, who they are, whether there's shoplifting; in a factory, how workers are doing; on construction sites, etc. This data is all very important nourishment for physical AI. Our first step is to put this data into a video data lake, let AI truly understand the real world through it, and turn these companies into future physical-AI-first companies.

Zhao Liang:

Feels like we could extend into world models and physical AI, but it's truly closely related. What you're doing is heavy-duty work.

EverMind's General-Purpose Memory Architecture and Business Model

Zhao Liang:

Actually, EverMind — Chen Tianqiao has always wanted to solve heavy-duty problems. EverMind may be similar to Memories in using brain-like architecture. My question for Ms. Han: is this architecture a general-purpose memory architecture, or is it aimed at a specific scenario? Regarding the business model for agent memory, what do you think it will look like?

Han Yunyun:

Mr. Shen just shared a lot of insights — we've known each other for a long time. He chose a path with great technical value, has an excellent reputation in Silicon Valley, and has received client validation. From EverMind's perspective, we happened to choose another path — trying to build general-purpose memory. This is very hard to define. As Mr. Shen said, every individual or enterprise building agents has their own understanding and needs for memory — a diverse range — abstracting that out is very difficult. But we happen to have the conditions to face this challenge early.

We do this in stages. The first stage abstracts the memory architecture from commonalities across several internal applications — health, future prediction, enterprise collaboration. We abstract the commonalities. Last year, when we open-sourced, we defined a four-layer structure, inspired by brain science. The bottom layer is the interface — like human senses. First, we define this interface for large amounts of data — context, localized data, heterogeneous data lake data, streaming data (data continuously flowing in from various applications) — and then give it to AI. Internally there's a memory layer that turns raw data into more structured memory, with CRUD and update mechanisms. Above that is an index layer for retrieval convenience. All memory ultimately needs to be recalled for agents; recall accuracy, latency, and cost are all things users feel in practice. So we've done a lot on memory recall efficiency, even building dedicated models and MSA (sparse attention). Above that is an agent layer, trying to handshake with the application layer. This layer mainly does two things: first, define the various ways and interfaces memory might be used in agent development and applications; second, leave ample customization space, actively exposing it, so others can use our capabilities in reverse through their own secondary development and customization. Through open source, we co-build use cases with many developers and clients, continuously improving a more general memory infrastructure.

On this architecture, we've achieved SOTA recall accuracy on established benchmarks, minimized latency and cost (1/10 of LLM full-context), and extended context to 100M through model-layer innovation. All this work has been shared through rigorous academic papers and open-source work, and we've received Oral Presentation invitations at this year's top industry conferences ACL and KDD.

That's the first step.

The second step is to provide more tangible value to clients and the developer community around key industries and scenarios. Just over a week ago, we released Raven — a complete memory-native harness. We found that just doing memory requires developers to do a lot of work to fully utilize it — they have to research various agent frameworks, integrate various skills and tools themselves, and other harness frameworks themselves limit memory capability. So we decided to build a better harness. We evaluated all skills on the market, including our own, and after testing and data comparison, selected nearly 100,000 skills for a skill hub, packaged into a harness together with memory. This harness is called Raven — named after the two ravens in Norse mythology that symbolize memory and mind. It surpassed 1,000 stars in about three days. We always do open source first, then commercial support. Raven already achieves code-level rewriting and can optimize invocation paths as usage grows, including skills that can be updated, even rewriting and updating parts of code.

Later we also built more to-C applications. We found some users are still new to AI and don't have strong coding ability, and running local agents is difficult. So we built a to-C application around Memory Hub called EverMe. Many capabilities will be integrated into it later. As Mr. Si mentioned, long-chain memory management — we're now integrating reasoning and memory for long chains. Memory itself supports long-chain task management, including memory isolation and sharing under multi-agent collaboration. But we found users mostly update memory as they reason, so we integrated the deep research agent — previously known in the industry as the "reasoning mini-cannon" — into Raven. Later, when people use Raven, it comes with built-in reasoning capability, and on top of that, memory is co-evolved with it. We've achieved continuous memory updates across 400-step reasoning. In such scenarios, when doing information-dense, long-duration tasks, it can stay continuously online without having to "retrain" from scratch mid-way. This is the process of Raven self-iterating as users iterate.

More memory-related actions will be released soon. When an enterprise or a heavy knowledge worker continuously hands work scenarios to agents, their knowledge needs to be deposited throughout the system, and there will continuously be more output that can be shared by others. They can share their entire agent capability with others, and we need to support such users to lead the entire agent usage experience forward. Toward this goal, we built an end-to-end self-evolving engine called context model. From the model layer to the agent interaction layer to the to-C layer, we've achieved an overall data flywheel. Users can not only build agents around memory within the entire system ecosystem, but also iterate agents during use, and further share their iterated agents with others, becoming commercializable agents.

MemorFit's Distinctive Architecture Design

Zhao Liang:

Thank you, Ms. Han. Next, Mr. Si — you also have an open-source spirit. Mr. Si's open-source project Yak has over 10,000 stars. You also launched the AI agent MemorFit. Let's discuss what distinctive architectural design MemorFit has to solve specific problems, and whether there are unique designs at the memory layer that differentiate it from traditional RAG architectures.

Si Hongxing:

The cybersecurity industry has its own特殊性. We're basically all offline-deployed. For critical infrastructure, no data is allowed to leave the network. In offline deployment, you can't use the strongest models — Claude, GPT-5.6, o1, o3 are definitely out. Including domestic DeepSeek V4 pro, Qwen 3.1/3.7 max — although open source, offline deployment for a team to use requires building a data center to drive such large models, which is unrealistic for a security team. So overall it's all offline deployment, which means most models used are 27B. Not that strong, so when executing long chains of tasks, they face a lot of spinning — continuously calling one tool, getting stuck and unable to get out. It can still call, but to complete the task it starts hallucinating, just keeps going. That's when you need a reflection and spin-detection mechanism. When such behavior is detected, you force a break. The first step after breaking is to search memory — see if similar problems have been encountered historically, how they were solved before, then reinject context and force reflection to break out and proceed. This is a big differentiator.

The second is memory generation during the process — working while generating memory, which is fine. But what's worth remembering versus just dumping the entire task in? We built an external observer model. When it falls into spin or deposits an experience, it extracts a JSON-format memory entry. This entry has a scoring mechanism — called the callback model — scoring from eight dimensions: relevance, timeliness, sentiment coding preference, user preference, practicality, and about seven or eight others. AI scores it. Some memories are old and expired, so when recalled they're pushed back, with a weight, and weighted scoring is applied at recall. Next execution dynamically injects this memory to complete the target task. This is one way MemorFit differs from traditional harnesses. Everyone can download and use it — it's full of cybersecurity attack and defense tools at MemorFit.ai.

In Education, What Should Be Remembered and What Should Be Faded?

Zhao Liang:

Thank you, Mr. Si. Mr. Si mentioned what should be recorded and what should be faded in agents. Education has the same problem. Teacher Zhang, we're worried that after students and teachers heavily use AI tools, it might have a reverse effect, ultimately limiting student development. Last year's teaching methods may not work this year. In education scenarios, regarding memory — what should be remembered, what should be faded, and who defines the boundaries? Please share on these specific issues.

Zhang Minsong:

I'll answer this from a different dimension. Education scenarios shouldn't be considered purely from a technical perspective; you also need to consider the fundamental logic of education. Education is a closed-loop process — learning is also a closed loop. With input comes feedback. Looking at in-school products, most schools now advocate and implement the concept of precision teaching, essentially data-driven integrated teaching-learning-assessment. By tracking, recording, and analyzing students' learning process data to make teaching decisions, this is a reform from "experience-driven" to "data-driven."

Most out-of-school learning products also follow the measurement-learning-practice-assessment learning loop. All of this involves evaluation — student evaluation is one of the most important indicators of education quality. With evaluation comes diagnosis; with diagnosis comes improvement. This closed loop also applies to learning-oriented education agents. Through this closed loop, the logic of memory updates and decoupling becomes very clear. Without this closed-loop data feedback, everything is presupposition. This is what I consider the core principle of Agent Memory design in education scenarios. For reference only.

Originally published by Unique Research on Unique Research Substack on August 4, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research