跳到正文
非凡资本

UNIQUE RESEARCH / ENGLISH ARTICLE

How Many Times Did You Press ENTER Today? An AI Company Is Watching That Signal

Original · Unique Research · 2026-04-20

Editor's note: The first-person report and its judgments belong to the original Chinese author. This English rendition retains the opening essay, full interview narrative, and all four Q&A turns. Product claims, technical pipeline descriptions, retention metrics, and market comparisons are source or speaker claims, not independently audited findings. Company, personal and product names are transliterated where official English forms remain unverified. The source is dated April 20, 2026.

Unique Awards

How Many Times Did You Press ENTER Today? An AI Company Is Watching That Signal

What is truly worth being understood by AI may not be what you said, but the moment at which you made a decision.

"

Every ENTER is the moment when a human says, 'I have decided.' Sending a message, confirming an operation, submitting a piece of code. This is the single frame with the highest information density in an entire workday, and the moment when a person's intent is clearest.

How many times did you press the ENTER key today?

You don't know. Nobody knows. But there is a product that is seriously counting it.

Their logic is this: every ENTER is the moment when a human says, "I have decided"—sending a message, confirming an operation, submitting a piece of code. This is the single frame with the highest information density in an entire workday, and the moment when a person's intent is clearest.

What AirJelly does is to quietly appear in those moments, remember what you did, and then gently nudge you when it is time to remind you.

Song Ziwen (Ziwen) is the co-founder of AirJelly. We talked for about an hour and a half. The sentence that impressed me most was not some grand narrative, but this:

"Our design team now uses AirJelly to modify AirJelly's own code."

A product using its own product to rebuild itself sounds like a tongue twister, but it is precisely the most honest proof an AI company can offer.

ChatGPT Is Powerful, but It Knows Nothing about Your Work

Let me start with a truth many people overlook.

Almost every AI tool on the market today shares a common fatal flaw: they don't know you.

ChatGPT is powerful, but every conversation starts from zero. Notion AI can only see what is inside Notion. Cursor only knows your current code repository.

You are simultaneously pushing forward three to five projects every day, jumping back and forth between Slack, Figma, email, Notion, and Terminal—no single AI knows what you are working on as a whole.

Ziwen calls this phenomenon "context fragmentation."

"You promised something in Slack, received a deadline change in email, and modified three files in code—this information is scattered across different tools, and your brain is the only glue."

But the human brain forgets, overlooks, and loses context when switching between tasks.

This is where AirJelly begins: not to build an AI that answers faster, but an AI that remembers what you were busy with yesterday.

It Is Not a Copilot, but the Kind of Companion That Is Always There

A Copilot comes when you call it. A Companion is always there.

The difference between these two words is much greater than it appears on the surface.

Ziwen explains: Copilot assumes that the user knows when to ask it. But in reality, at the moments when help is truly needed—a forgotten promise, a deadline about to expire, a work thread that was interrupted and never returned to—the user often doesn't even realize it.

"Passive AI has a fatal blind spot: it waits for you to speak up. But the things you've forgotten, you won't ask about."

AirJelly's approach is the exact opposite: it continuously perceives what you are doing through the screen, and judges for itself when to intervene and when to stay quiet.

Technically, its workflow roughly looks like this:

Scheduled screenshot → VLM understands screen content → triggers memory retrieval → episodic memory + semantic memory → vector matching with task library → LLM judges whether to create a new task

Each screenshot frame is not random, but precisely triggered at the moment the user presses ENTER—the moment when the decision is clearest and information density is highest.

"This way, when you are busy, it screenshots frequently; when you are zoning out, it doesn't screenshot; every frame it captures is the state after 'an action has been completed,' not the intermediate state while you are halfway through typing."

How Does Proactive Reminding Achieve "Useful but Not Annoying"?

Almost every product that does proactive notifications ultimately dies from the same problem: too much interruption.

To solve this, AirJelly has split its notification system into four layers:

Signal detection → Signal accumulation → Scene matching → Push decision

It does not push upon seeing a single signal. The signal must first accumulate to a threshold, then match a known scene (such as "unfulfilled promise" or "high-priority task left untouched for too long"), and finally pass through a push-decision layer: cooldown period, global frequency limit, and task-level anti-duplication.

I asked Ziwen how this boundary is defined.

He paused for a second, then said:

"The boundary between genuinely effective proactive support and random reminding is whether it knows you need it before you do. If the user has already remembered, and then it reminds them, that is interruption. If the user hasn't remembered yet, and it arrives, that is value."

Event Is Fact, Task Is Understanding

Many people's first concern about AirJelly is: will it capture the wrong things?

Ziwen breaks this question down in detail.

Inside AirJelly there are two levels of concepts: Event and Task.

An Event is a raw observation: "The user discussed an API design proposal with Zhang San in Slack." This is always recorded—where there is a screenshot, there is a record.

A Task is a goal-level abstraction: "Complete the API interface refactoring." This is only created when the LLM judges that the event represents an actionable work goal.

More critically, task creation is not simple keyword extraction. The LLM performs a "dimensionality elevation" action: abstracting a specific event into a work goal that can accommodate more related events in the future.

For example, when you discuss an API proposal with someone, it doesn't just record a memo about the API proposal—it understands that this is part of the "refactor the interface" task, and groups subsequent related screenshots into it.

"It's not capturing keywords; it's reading the context."

Truly Usable Memory Must Satisfy One Condition

Almost every major model company is building an AI Memory feature. But anyone who has used one knows that most of the time it feels unreliable.

Ziwen has his own judgment on this: traditional search is for when you know what you are looking for; AI memory solves the kind of need where you don't know what words to search with.

"Who did I discuss that API proposal with last week?"

You actually don't remember what it was called, don't know which tool it was in, and can't recall the specific date—but AirJelly knows. You ask directly, and it replies instantly: last Wednesday in Slack, you and Li Ming discussed API gateway, and he suggested using Kong instead of building in-house.

This is the essential difference between Contextual Memory Recall and traditional search: its indexing dimension is your work context, not keywords.

They call this part Full Traceability—every memory can be traced to its source: which screenshot, at which point in time it was extracted.

"Traceability is especially important for proactive Agents. Because you need to be able to verify that it hasn't misunderstood you. If it says something and you don't know where it came from, trust breaks."

How Far Does Privacy Need to Go Before Users Dare to Use It?

What AirJelly does—continuous screenshots, understanding your work activity, extracting memories—if the data isn't local, most people simply won't touch it.

Ziwen says they have an internal principle: the capability boundary must be smaller than the user's trust boundary.

This means that even if technically more can be done, if it exceeds the user's psychological comfort zone, it is not done.

All activity records, memories, and tasks are stored locally (SQLite + LanceDB). AI model calls go through their own AI Gateway, but only the text fragments needed for the current processing are sent to the model—not the entire user dataset. No cloud backup of memories, no cross-device sync of memories—unless the user actively chooses to.

"Because these capabilities touch the red line of trust. Users aren't afraid of AI seeing; they're afraid of AI talking nonsense, spreading indiscriminately, and misusing data."

The Moat Is Not in Features, but in Six Months of Your Work Memory

I finally asked Ziwen a rather direct question: with large model capabilities advancing so quickly, where exactly is a startup's moat?

His answer had two layers.

Short term: engineering capability. VLM screenshot understanding, memory extraction, task attribution, proactive pushing—making each link accurate and reliable requires a great deal of refinement, not something a big company can casually do with a general model.

Long term: your work memory itself.

"When AirJelly has accumulated six months of your work memory, the switching cost is not at the feature level, but at the context level. You switch to another tool, and your work memory is gone."

Big companies build general platforms; startups build deep vertical scenarios. The last thing to compete on is model capability; the most important thing to defend is "deep understanding of the user's work process."

Ziwen believes the current stage is more like "when PCs first got graphical interfaces"—the direction is set, but the interaction paradigm hasn't been finalized. Should an AI Companion be a floating avatar? A sidebar? Or an invisible background process? These are still being explored.

But there is one thing he is very certain about:

"AI will shift from 'you go to it' to 'it is by your side.' Just as everyone has a browser on their desktop today, in the future everyone will have a continuously running AI Companion on their desktop."

Selected Interview Q&A

Q1: How does AirJelly turn signals scattered across different tools into a continuous workflow, rather than a pile of disconnected events?

Song Ziwen (AirJelly co-founder): The technical pipeline is: scheduled screenshot → VLM understands screen content → triggers memory retrieval → episodic memory + semantic memory, then each event also goes through task association—finding candidate tasks through vector search, then using the LLM to judge whether it should be assigned to an existing task or create a new one. This way, fragmented screenshots become a continuous workflow narrative. The hardest engineering challenges are: noise filtering in VLM understanding, context continuity across windows, and the accuracy of task attribution.

Q2: Why can the ENTER key become the key entry point for understanding user intent?

Song Ziwen: ENTER is the most special—it almost always represents "an intent has been completed": sending a message, submitting a search, confirming an operation, executing code. By triggering screenshots at this moment, what we capture is not the intermediate state of halfway typing, but the final state after the user has made a decision. This brings two distinct experiences: screenshot frequency adapts to user activity (you get more screenshots when busy, none when zoning out); and every frame contains high-density information of a "completed action."

Q3: What retention metric do you value most, and how do you judge that a user has truly "gotten it running"?

Song Ziwen: What we value most is not daily active users, but the frequency with which users ask AirJelly about their past work—that is, the number of times the action of "retrieving a piece of information I forgot" occurs. The more this happens, the more dependent the user is on its memory. Once someone starts asking "who did I discuss what with last week" instead of digging through Slack and email themselves—the product has truly started running for them. The sign of shifting from "novelty" to "can't live without" is that search-behavior migration has occurred.

Q4: Among the five variables of multimodality, Agent, memory, on-device, and proactivity, which will become the decisive variable first?

Song Ziwen: I believe it is memory. Because it is the foundation of all other capabilities—an Agent without memory starts from zero every time; proactivity without memory is just guessing; multimodality without memory is one-time understanding. Memory is the turning point that transforms AI from a "powerful tool" into a "trustworthy companion."

Typesetting version: Unique Research-A typesetting (full original text)

Originally published by Unique Research on Unique Research Substack on April 20, 2026. This page preserves the public article for reading on UniqueCapital.

View the original publication ↗
← Back to English research