I.
For fifteen years I've been building social platforms.
I started MaiChang in 2013 — a music-interest social network in China that eventually grew to fifty million users. Before that, I spent years at Huawei leading voice algorithm research. I filed nine patents along the way — one of them a U.S. patent on double-talk detection, and one an international PCT patent. Most of them I don't think about anymore.
What I do think about, almost every day, is this: in fifteen years of building systems where humans talk to each other through software, I never built a system that understood what was actually being said.
Not the words. The relationships.
Every platform I've built has been a place where connections form and dissolve in patterns I could measure but never model. Who talked to whom. How often. For how long. I could graph it. I couldn't explain it.
For most of those years, nobody could. You couldn't put "the shape of a friendship" into a database schema.
And then, very quietly, that started to change.
II.
Here is a truth that I think most people in AI haven't sat with long enough:
The field has been scaling models for two years. It has barely started scaling memory. And it has not yet begun to scale society.
This is the order of operations for AI, whether we plan it that way or not.
The first wave was about making models bigger and more capable. GPT-3, GPT-4, Claude, Gemini, DeepSeek. That wave is largely won. The frontier labs know how to build bigger models, and the gap between them is narrowing every quarter.
The second wave — the one we're in now — is about giving those models persistent memory. Letta, Zep, and a handful of others have started building this layer. It's the right instinct. But most of what exists today is memory as storage: a vector database with a retrieval function bolted on. You remember what I said last Tuesday. You do not remember who I am.
The third wave, the one almost nobody is building yet, is about AI society. It's about agents that exist in relation to other agents, that develop strategy, trust, deception, cooperation — the things humans call social intelligence and philosophers call theory of mind. It's about AI that doesn't just speak fluently, but speaks with someone, over time, with something resembling a history.
This is where I think Sonari lives.
Not in the model layer. Not even in the memory-as-storage layer. In the layer above both: the place where memory and agents become a living system.
I don't think of this as a product roadmap. I think of it as the next frontier of a field that, surprisingly, hasn't noticed it's there.
III.
Let me tell you what convinced me.
My team and I have been operating voice-first social products internationally since 2019. The first one was sigo — a voice social product in the Middle East, launched in the same era as Yalla. We eventually took it down for other reasons, but what sigo taught us about voice-first user behavior in emerging markets became the foundation for everything that came next.
Today, that portfolio includes:
- peplive — an AI voice agent social network with multi-language coverage and AI Friends in production for over three years
- pepstar — a multi-agent social games platform that publicly markets itself as an "AI Agent Mesh"
- veco and uho — voice social verticals serving Southeast Asia and the overseas Chinese diaspora
Across them, we serve users in 4 emerging markets — roughly 100,000 combined monthly actives across the portfolio. Small by social platform standards. They were meant to be.
I'll be honest about something: these apps were never going to scale like MaiChang did. I knew that when I started them. The voice social category in emerging markets has thin margins, intense competition, and a user base that churns hard. I chose them anyway. Not to build the next Yalla. To build a research environment.
What you see on peplive.com and pepstar.tw — "AI Voice Agent Social Network" and "AI Agent Mesh" — is Sonari, in production, in public, today. Not a roadmap. Not a pitch. Already running.
And here's what running it for years taught me, the lesson that turned this from a portfolio of apps into a company called Sonari:
Every time we added AI companions to these products, users loved them. For about three sessions.
Then they would stop coming back to the AI specifically. They would still come back to the apps — to talk to other humans — but the AI, the thing I'd been so excited about, would quietly get ignored.
I watched this happen across multiple experiments. Different models, different personas, different voices, different onboarding flows. The failure mode was always the same. And the user feedback, when we asked, was always the same, phrased in a hundred different ways:
"It's nice, but every time I come back it feels like meeting a stranger."
I'm a founder, not a researcher, but I've been around long enough to recognize when a failure mode is trying to tell you something. The problem wasn't that our AI wasn't smart. Some of the models we tested were extraordinarily smart. The problem was that the AI had no continuity. It couldn't remember. It couldn't grow with the user. It couldn't build a relationship.
The missing layer wasn't a bigger model. It wasn't even a better persona. It was memory that captured a relationship, not just a conversation.
That's when Sonari stopped being an experiment and started being a company.
IV.
The insight I care about most, the one I'd argue for in any room, is this:
Human memory isn't a database. It's a living structure. And AI memory, if we want it to actually work, has to become one too.
When you remember a friend, you don't retrieve a log of your last twelve conversations. You remember a feeling. A trust level. An unspoken rhythm of jokes and references that you've built together over years. You remember what topics to avoid and which ones light them up. You remember the shape of the relationship, and that shape itself is what makes your interactions feel continuous.
No current AI system has this. Not Character.AI, which remembers what you said. Not Replika, which remembers who you are. None of them remember the relationship between you and the AI — because the category of thing we need to store there doesn't exist yet in any production system.
We've been building that category at Sonari. We call it Relational Memory, and it sits at the top of a five-layer stack. The lower four layers — working, episodic, semantic, procedural — are well understood in cognitive science and have been discussed across the AI memory community. What no one has built is the fifth: a quantified, production-validated model of relationship state. And what no one has done is take all five layers from theory to production in real voice social applications. We've done both.
I won't pretend we've solved it. We have an architecture, a data model, a production deployment serving our own apps, and a growing pile of real-world observations about what works and what doesn't. We're early. But I believe more strongly every week that this is the right place to stand.
Character.AI remembers what you said.
Replika remembers who you are.
Sonari remembers the relationship between you and your AI.
That one line is the closest I can get to what we're building, in one breath.
V.
There's a second thing I care about, and this one is less obvious.
Social intelligence cannot be trained in isolation. It has to emerge from society.
Every AI companion product I've seen, including my own early attempts, tries to make a single AI "more social" by training it on more human-labeled dialogue. This will get you somewhere. It won't get you to an AI that actually understands social situations, because social understanding is not a property you can pour into a single mind. It's a property that emerges between minds, through repeated interaction, cooperation, deception, and repair.
The academic consensus on this isn't new — Meta's Cicero work on Diplomacy, DeepMind's multi-agent cooperation research, Sakana AI's work on evolutionary agent systems, and growing academic interest in Werewolf and Avalon as theory-of-mind benchmarks. Researchers have been saying for years that social reasoning games are the theory-of-mind equivalent of AlphaGo's Go: deceptively simple on the surface, computationally brutal once you start requiring actual understanding of other minds.
What I notice is that almost all of this work lives in research labs. Very little of it has been plugged into production systems with real users, real stakes, and real social dynamics.
We recently turned on something inside Sonari that I've been thinking about for months. We built an autonomous community of one hundred AI personas and let them start playing social reasoning games against each other — continuously, 24 hours a day, seven days a week, no human in the loop. They argue. They form teams. They get fooled. They learn. We capture the data and feed it back into the system.
It's small. It's early. It's not going to beat GPT-5 at anything tomorrow. But every hour it runs, it generates training data that no human-labeled dataset can produce — because the data is about agents making real social decisions under uncertainty, competition, and cooperation, against other agents that are themselves learning.
I think this is going to compound in a way that, from the outside, will look very sudden.
VI.
I want to say something honest about timing.
I have been building in voice and social for long enough to see multiple waves come and go. I watched the early mobile social era. I watched the live streaming era. I watched the voice chat era. I watched the AI companion era that crashed into Character.AI's ceiling. Every wave has a moment where the underlying conditions shift and the thing that was "too early yesterday" becomes "obvious tomorrow."
For AI memory and AI society, that moment is now.
Real-time voice AI just matured. GPT-4o Realtime, Claude Voice, MiniMax Speech — the pieces that make voice-first AI personas feasible arrived within the last eighteen months. Memory as an infrastructure category has been validated by the success of Letta, Zep, and the broader AI memory conversation. Character.AI's founders returned to Google in 2024, leaving the companion category in transition. Emerging markets in Southeast Asia, the Middle East, and Africa are going voice-first faster than any Western analyst seems to realize. And the academic consensus on multi-agent self-play as a route to social intelligence has quietly solidified over the past year.
Every precondition for Sonari to exist became true within the last twenty-four months. None of them were true when I started MaiChang in 2013. Some of them weren't even true when I sketched the first version of Sonari's architecture last year.
This is the moment. I don't say that as marketing. I say it because I've been waiting a long time to feel this way about a problem.
VII.
Here is what Sonari is, stated as clearly as I can:
We are building the agent society infrastructure for voice-first platforms. That has three concrete pieces.
A Memory Layer with five kinds of memory — working, episodic, semantic, procedural, and relational — built as production infrastructure, not as a research artifact. Deployed in our own apps today. Being prepared to open to developers later this year.
An Agent Runtime that handles persona registration, session management, cross-app routing, and all the unglamorous work of making AI personas actually run in production. We've been hardening this in our own apps, iterating on it through real production usage. It isn't pretty. It works.
An Agent Society — our live, 24/7 autonomous community of AI personas that play social reasoning games, generate self-play training data, and evolve over time. It is running as I write this. It is going to get much bigger.
The goal is not to build another AI companion app. The goal is to build the infrastructure layer that every voice-first AI companion app will need, whether it's ours or someone else's. Our own apps are Sonari's first customer. They were always meant to be.
I'm not going to pretend this isn't also a business. We're raising a Series A to open Sonari's platform to developers, deepen the Memory Layer research, and build the agent society into something that can serve the whole industry. I'll say more about that separately.
But the reason I'm writing this essay, and not a pitch deck, is that I think the intellectual frame matters more than the funding frame. If I'm right about what the next frontier of AI looks like, the funding follows. If I'm wrong, no amount of pitching fixes it. So I'd rather tell you what I actually believe and let you decide if it's worth your time.
VIII.
A closing thought.
For fifteen years I built systems where humans talked through software. I watched relationships form inside those systems and I watched them break. I watched what happened when a product understood its users deeply, and I watched what happened when it didn't. The difference between those two outcomes is almost always the difference between a company that grows and a company that quietly dies.
AI persona products are about to learn this lesson the hard way. Most of them are still optimizing for the quality of a single response. The ones that win are going to be the ones that optimize for the quality of a relationship, measured across months and years. Measured across how the AI grows with the user. Measured across whether it remembers.
If you're building something that needs persistent memory, voice-first interaction, or agents that can actually understand social situations — I want to talk. If you're a researcher working on theory of mind, multi-agent self-play, or the long-horizon memory problem — I want to talk. If you're a developer tired of AI companions that forget you every session — come try what we're building.
And if you're the kind of person who reads a 3,500-word essay about AI memory on a Tuesday afternoon and thinks "yes, this is the next frontier, and I want to help make it real" — then I think we already agree on enough to have a conversation.
Sonari is at the beginning. But we are genuinely, practically, in-production at the beginning — not the powerpoint beginning, and not the "we'll figure it out after funding" beginning. We've been operating voice social apps internationally since 2019, deep into the agent society experiment that I think is going to change everything, and one Manifesto into telling the world what we're doing.
If you see what I see, let's talk.
If you see what I see — let's talk.
Email Ronnie →