[{"content":" TL;DR: AI application architecture will be divided again according to responsibility. Cognitive structures that help models understand, reason, plan, remember, and correct themselves will gradually be internalized by stronger models and runtimes. Structures that define where goals come from, how systems connect to reality, who may act, what constraints apply, and how outcomes are proven cannot be self-granted by a model—and will not disappear. Architecture will shift from compensating for model weaknesses to establishing trustworthy real-world boundaries for powerful actors.\nIntroduction: AI architecture will not simply become thinner; its structure will migrate When discussing the future of AI application architecture, people often move between two claims. One says that as models become stronger, prompts, RAG, workflows, agent frameworks, memory, and orchestration layers will all be absorbed, leaving only a model. The other says that stronger AI will make peripheral systems more complex.\nNeither claim is complete. AI application architecture will neither vanish as a whole nor expand indiscriminately. Its responsibilities will be redistributed: some external structures will lose their independent necessity as models improve, while others will become more important precisely because models gain the ability to act.\nThe key distinction is not whether a technology sits inside or outside a model, nor whether it is currently called an agent, platform, or middleware. It is the question it answers:\nDoes it help a model answer how to complete a task? Or does it specify which reality the AI relies on, whom it represents, what it may do, what it must obey, and how it proves its result? The former is a cognitive-compensation structure; the latter is a real-world action structure. The long-term evolution of AI architecture is a process in which cognitive compensation is continually internalized while real-world action structures are retained and strengthened.\n1. The real dividing line: compensating cognition or defining real-world boundaries The future of a structure does not depend on whether it is called Prompt, Context Engineering, RAG, Memory, Skill, Tool Calling, MCP, Workflow, Agent, Multi-Agent, A2A, Guardrails, Evaluation, or Observability. It depends on the responsibility it carries.\n1. Cognitive-compensation structures answer: “Can the AI do it?” These structures help a model understand requirements, decompose work, select tools, manage context, discover errors, and organize output. They compensate for cognitive capabilities the model cannot yet perform reliably.\n2. Real-world action structures answer: “May the AI do it, should it do it, and is the result valid?” These structures define goals, authority, access to systems and data, policy constraints, approval procedures, audit trails, accountability, and proof of completed actions. They do not merely make a model more capable; they make actions in the real world legitimate, controllable, and verifiable.\nThis distinction matters because improved model intelligence can replace many cognitive methods, but intelligence alone cannot create authority, ownership, institutional rules, or evidence.\n2. First trend: external structures that compensate for model cognition will be internalized 1. Why cognitive-compensation structures appeared Prompt templates, retrieval pipelines, planning workflows, agent loops, memories, tool-selection rules, and evaluation chains emerged because earlier models were not sufficiently stable at understanding context, planning, remembering, or self-correcting.\n2. Why cognitive methods will be absorbed upstream Once a method becomes broadly reusable and can be learned from data or optimized at the runtime layer, it tends to move closer to the model. Better reasoning, longer context, native tool use, multimodal understanding, persistent memory, and more capable runtimes all reduce the need for application teams to reconstruct the same cognitive scaffolding repeatedly.\n3. Internalization does not mean immediate disappearance Specific technologies will remain useful for a long time. RAG will still matter where knowledge must be current, private, or attributable; workflows will remain useful where processes require determinism. But their role will shift from compensating for general intelligence toward providing domain-specific data, reliability, and operational control.\n3. Second trend: structures that define the relationship with reality, action eligibility, behavioral constraints, and proof of results will not disappear A model cannot determine on its own whose interests it represents, which account it may operate, which system it may access, how much it may spend, whether a human approval is required, or what evidence counts as a completed action.\nThese issues belong to organizations, law, systems of record, contracts, permissions, and accountability. They therefore remain outside the model, even when the model becomes much better at reasoning and execution.\n1. Goal sources cannot be invented by the model Models can help formulate goals, but legitimate goals must originate from people, organizations, contracts, policies, or explicit delegation.\n2. Connections to reality require stable interfaces Enterprise systems, databases, payment systems, devices, identity systems, and external services need stable protocols, permissions, and failure handling. This is not merely a cognitive problem.\n3. Action authority must be explicit and revocable The more powerful an agent becomes, the more its access, scope, spending limits, and emergency stops must be explicit. Delegation needs boundaries and revocation mechanisms.\n4. Constraints and accountability cannot be replaced by a better prompt Policies, compliance requirements, approval workflows, audit records, and responsibility chains cannot rely on a model’s self-description. They require independent enforcement and evidence.\n4. The architectural center of gravity will move: from “making models smarter” to “making actions trustworthy” Early AI application work naturally focused on improving model performance: better prompts, retrieval, planners, memories, and workflows. As those capabilities become increasingly native, differentiation will move toward trustworthy action in real settings.\nFuture architecture will pay more attention to:\nidentity, authorization, delegation, and revocation; access to systems of record and external tools; policy enforcement, approvals, and risk limits; traceability, evaluation, and evidence of outcomes; human intervention, rollback, and incident handling; durable context that belongs to users and organizations rather than to a single model provider. The question will change from “How can we make the model finish this task?” to “Under what authority, within what boundary, and with what proof can the system complete this task?”\n5. What this means for agents, platforms, and application teams Agents will not disappear, but they will become less valuable as thin wrappers around model cognition. Their durable value lies in serving as governed execution systems that bind goals, context, skills, tools, permissions, operations, and evidence.\nPlatforms will likewise be judged less by how many prompt chains they offer and more by whether they can provide reliable identity, data connections, policy controls, auditability, observability, and interoperability.\nFor application teams, the competitive advantage will increasingly come from understanding real workflows and embedding AI into the institutional and operational constraints of a specific domain.\nConclusion: methods enter the model; boundaries remain in reality AI application architecture is not moving toward “only the model.” It is moving toward a clearer division of labor. Methods that mainly compensate for the model’s temporary cognitive limitations will gradually be absorbed into models and runtimes. Structures that define real-world goals, permissions, constraints, accountability, and proof will remain external—and become more important.\nThe long-term architectural task is therefore not to build ever more elaborate scaffolding around a model. It is to create reliable boundaries through which a powerful AI can understand reality, act in reality, and remain answerable to reality.\n","permalink":"/en/posts/ai-application-architecture-evolution/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: AI application architecture will be divided again according to responsibility. Cognitive structures that help models understand, reason, plan, remember, and correct themselves will gradually be internalized by stronger models and runtimes. Structures that define where goals come from, how systems connect to reality, who may act, what constraints apply, and how outcomes are proven cannot be self-granted by a model—and will not disappear. Architecture will shift from compensating for model weaknesses to establishing trustworthy real-world boundaries for powerful actors.\u003c/p\u003e","title":"The Evolution of AI Application Architecture: Methods Move into the Model, Boundaries Stay in Reality"},{"content":" TL;DR AI flattens old gaps while creating new layers of inequality. That is not a bug but the default path of technological change. Printing needed public education; the internet needed open-source institutions. AI’s equalizing effects will not complete themselves. The decisive arena is not the laboratory, but institutional design.\nThis series has traced four journeys. First, AI lowers the threshold for basic skills while widening gaps in higher-order capability. Second, credentials lose signaling power while judgment earned through experience gains value. Third, outsourcing thought can quietly erode judgment. Fourth, applications become more universal while control over AI concentrates.\nTogether, these arguments reveal a single pattern: technology creates the appearance of equality while producing new structures of stratification. Each layer of equalization is real—skills become easier to access, knowledge becomes more available, and tools become more widespread. But new differentiation and screening soon follow. Equality has never been a task technology can complete on its own.\nEvery Technological Equalization Needs Institutional Equalization The printing press lowered the cost of books and released knowledge from the monopoly of churches and aristocrats. Yet it did not automatically democratize knowledge. The first beneficiaries were already-literate merchants and artisans, while many rural and urban poor remained excluded for centuries. Publishing also created new elites who controlled production and selection.\nPublic education, literacy campaigns, public libraries, and copyright institutions eventually turned the press’s potential into broader equality. Technology created the possibility; institutions made it real.\nThe internet followed the same path. It promised information democracy, but research on MOOCs found that those already well educated benefited the most. The internet’s openness depended on later institutional innovations: the open-source movement, net-neutrality principles, Creative Commons, and public investment in digital infrastructure.\nAI now stands at the same point in this rhythm. Technical equalization is already well underway; institutional equalization has barely begun.\nWhy AI Makes This More Urgent AI differs from previous technologies in three ways.\nSpeed. The printing press diffused over decades, the internet over roughly fifteen years, and ChatGPT reached one hundred million users in two months. The window for social adaptation and institutional adjustment is dramatically shorter.\nScope. Printing replicated the carrier of knowledge; the internet replicated its distribution. AI replicates parts of knowledge processing itself: reasoning, generation, and judgment. Its impact is therefore deeper and broader.\nConcentration. Printing and early internet infrastructure were relatively distributed. Frontier AI is concentrated in a small number of companies because training a leading model requires hundreds of millions of dollars. That makes democratization more dependent on corporate choices and policy constraints.\nThese differences mean AI’s stratifying effects may arrive faster, run deeper, and be harder to reverse. Without timely institutional action, new layers can solidify within only a few technical cycles.\nStratification Is Happening, but Its Direction Is Not Fixed AI-era intellectual stratification is unfolding on three levels.\nFirst: access to tools—consumer equality versus producer control. Nearly everyone may use AI, while the power to create it concentrates. The institutional questions are clear: should concentrated AI markets face antitrust intervention? Should foundation models be regulated as public infrastructure? Should data contributors share in the gains?\nSecond: distribution of capability—old skills flatten while new skills diverge. Coding, writing, and translation become easier, while judgment, metacognition, aesthetic sense, and complex decision-making become more valuable. These capabilities are distributed more unevenly and depend more heavily on practice resources and high-quality educational environments.\nThird: cognitive structure—different ways of thinking. Under similar social conditions, different patterns of AI use can produce different cognitive results. Those who use AI as a tool while retaining independent judgment may diverge from those who treat it as a replacement for thinking.\nThese three layers are dynamic, and each one reinforces the next.\nFour Pieces of the Rule-Making Puzzle Institutional innovation has four urgent directions.\nUniversal AI literacy. This is not merely programming education. It means understanding what AI can and cannot do, using it structurally—think first, verify, then integrate—and practicing independent reasoning without AI.\nAlgorithmic transparency. When AI materially affects a person through loan decisions, job screening, or ranked results, that person should know the basis of the decision, how their data is used, and how to correct or appeal it.\nRedistribution of data benefits. People who contribute data should not indefinitely provide free training material while model owners capture all gains. Possible mechanisms include data taxes, benefit-sharing arrangements, and public-model funds.\nPublic compute and open ecosystems. Governments can lower entry barriers by investing in public compute infrastructure and supporting open-model ecosystems. Like public libraries in the age of print, public compute can prevent a few firms from locking the market.\nThese pieces work together: literacy addresses cognition, transparency addresses trust, redistribution addresses fairness, and public compute addresses concentration.\nThink of It Like Building Highways AI equality resembles building highways. Roads make everyone move faster, just as AI helps everyone obtain information, generate content, and complete tasks more quickly. That is real equalization.\nBut highways also create new inequality. People living near interchanges benefit more; remote regions may wait decades for connections; some can afford vehicles and operating costs while others cannot. Likewise, people with resources and strong educational backgrounds are better positioned to turn AI access into real gains.\nHighways need traffic rules, driver licensing, public funds, and pricing adjustments. AI needs algorithmic transparency, AI literacy, public investment, and benefit redistribution. Without these institutions, infrastructure may amplify rather than reduce inequality.\nThe End Point of Equality Is Not in Technology The answer to whether AI narrows or widens intellectual gaps is neither simply one nor the other. AI does both: it flattens old differences and creates new strata. These are two sides of the same coin.\nIntellectual equality in the AI era is real, partial, and conditional. It is real because AI compresses the distribution of capability in some dimensions. It is partial because this occurs mostly at the tool-use layer, not at the layers of cognitive depth or power. It is conditional because its final effects depend on institutional design.\nTechnology-driven equality has limits. It removes barriers that can be encoded and automated, while creating new barriers that require judgment, experience, and power to cross. What determines whether equality can truly emerge is not how capable the next model is, but whether we build institutions that distribute AI’s benefits more broadly and allow everyone to be not only consumers, but participants and beneficiaries.\nThis is not an argument against technology. It is an argument for institutional innovation.\nAI will not bring equality by itself. People will.\n","permalink":"/en/posts/ai-equality-institutional-rules/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\nAI flattens old gaps while creating new layers of inequality. That is not a bug but the default path of technological change. Printing needed public education; the internet needed open-source institutions. AI’s equalizing effects will not complete themselves. The decisive arena is not the laboratory, but institutional design.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003chr\u003e\n\u003cp\u003eThis series has traced four journeys. First, AI lowers the threshold for basic skills while widening gaps in higher-order capability. Second, credentials lose signaling power while judgment earned through experience gains value. Third, outsourcing thought can quietly erode judgment. Fourth, applications become more universal while control over AI concentrates.\u003c/p\u003e","title":"AI and Intellectual Equality (5): Technology Creates Stratification, Equality Requires Rules"},{"content":" TL;DR AI applications are spreading rapidly: ChatGPT reached 100 million users in two months, Copilot is embedded in editors, and image models let non-designers create visual work. At the same time, the power to build and govern AI is concentrating in a few companies. Consumers gain unprecedented access, while producers gain unprecedented control. You receive the freedom to use AI, but surrender much of the freedom to define it.\nA young person in rural Africa can now use a phone to receive near world-class tutoring from a GPT-4-level model. In use, the experience can resemble that of a Silicon Valley engineer: the same interface and the same model. From the consumer side, this is an extraordinary democratization of knowledge.\nFrom the supply side, the picture is different. The model may be developed by a U.S. company, trained on infrastructure worth hundreds of billions of dollars, fed with global internet data, and aligned according to choices made by a small group of engineers. Users can access the system, but they do not decide how it updates, what data it uses, or what values it adopts.\nAI creates a historic split: use becomes more equal while control becomes more concentrated. This is not merely temporary; it follows from the industry’s technical structure.\nAccess Is Real, but It Is Not the Whole Story At the tool layer, AI genuinely lowers capability barriers. People who never learned programming can produce runnable code; people without design training can generate professional-looking images; people without data-science backgrounds can ask AI to perform complex analyses. Skills that once required years of training can become on-demand tools.\nThis broad access is meaningful. But being able to use a system is not the same as being able to control it.\nModel Costs Build a Wall Why does AI power concentrate? Start with the cost of building frontier models. Training requires massive compute, vast datasets, specialized talent, and extensive infrastructure. Stanford HAI’s AI Index shows that frontier research is concentrated among a small number of firms in the United States and China.\nThis differs sharply from the early internet. A student in a dormitory could create an internet startup; today, a student can build an API-based application but cannot independently train a frontier model from scratch. The difference is structural, not simply a matter of effort.\nAs a result, much of the world becomes a consumer of AI rather than a participant in its production. Global market growth is driven primarily by a small number of countries and companies.\nData Colonialism Is Underestimated Couldry and Mejias describe data colonialism as the transformation of human life into a continuing raw-material source for capital accumulation. In AI, that pattern appears in two ways.\nFirst, data contribution is disconnected from value distribution. Billions of users contribute searches, conversations, and feedback that help models improve, while most gains flow to companies that own the models. Second, the physical infrastructure of AI—data centers, GPU clusters, and high-speed networks—is geographically concentrated.\nThe digital divide therefore persists beneath the apparent universality of AI applications. Closing it requires public and global-scale investment, not merely free online courses.\nRegulation Cannot Keep Pace with Iteration The EU AI Act is the most comprehensive attempt to address these issues. It includes support for small firms, regulatory sandboxes, AI-literacy obligations, and restrictions on some uses of AI in unequal power relationships.\nIts core insight is important: AI does not automatically create equality; institutions can guide it toward more equal outcomes. Yet legislation moves in years while AI capabilities change in months. Compliance costs can also favor large firms that have the resources to absorb them, potentially widening the gap the rules seek to close.\nThe User’s Illusion: Freedom of Use, Not Power AI users have the freedom to use tools for writing, coding, analysis, and design. But they rarely participate in defining the system’s direction: training data, value alignment, privacy boundaries, updates, or shutdowns. Those decisions are made by a limited number of actors because entry costs are so high.\nAs Cathy O’Neil argues, mathematical models are not neutral; they encode choices and existing power structures. When designers are concentrated in a few companies and countries while users span the world’s cultures and social conditions, bias becomes not only a technical question but a political one.\nThe risk is especially serious when AI becomes infrastructure for search, recommendation, hiring, credit, and education. If these systems are controlled by only a few companies, users may have nowhere else to go.\nThink of It Like an Operating System AI’s distribution resembles the smartphone era. At the application layer, smartphones created real capability equality: people across the world can use the same messaging, video, and map applications. Yet the underlying operating systems determine app-store rules, update schedules, data collection, and privacy policy.\nAI may reproduce this structure. Users move freely among applications, while the underlying AI operating system—foundation models, training frameworks, and compute infrastructure—remains concentrated in a few firms. Once such an ecosystem becomes dominant, replacement is extremely costly.\nWhich Force Breaks the Balance First? Technology history shows that distributional outcomes are not predetermined. Acemoglu and Johnson argue that technology opens a space of possibilities, while institutions, political struggle, and policy decide who benefits. Early industrial growth did not automatically raise workers’ living standards; institutional change eventually altered the distribution.\nAI is similar. Public compute infrastructure, data-benefit redistribution, open-model disclosure, and active antitrust enforcement could restrain concentration. Without such interventions, AI may become more concentrated than mobile platforms because frontier-model training leaves so few viable players.\nAI access is real, and AI power concentration is real. The long-term balance depends on whether institutions can imagine and build rules that prevent power from naturally pooling in an era with unprecedented technical barriers.\n","permalink":"/en/posts/ai-access-power-concentration/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\nAI applications are spreading rapidly: ChatGPT reached 100 million users in two months, Copilot is embedded in editors, and image models let non-designers create visual work. At the same time, the power to build and govern AI is concentrating in a few companies. Consumers gain unprecedented access, while producers gain unprecedented control. You receive the freedom to use AI, but surrender much of the freedom to define it.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003chr\u003e\n\u003cp\u003eA young person in rural Africa can now use a phone to receive near world-class tutoring from a GPT-4-level model. In use, the experience can resemble that of a Silicon Valley engineer: the same interface and the same model. From the consumer side, this is an extraordinary democratization of knowledge.\u003c/p\u003e","title":"AI and Intellectual Equality (4): Access Becomes Universal, Power Concentrates"},{"content":" TL;DR Academic credentials are losing their power as signals of capability because AI can master textbook knowledge. The premium is moving elsewhere: toward judgment earned through real responsibility. When knowing something is no longer scarce, what you have actually done becomes the new screening criterion.\nTwo résumés sit before a recruiter. Candidate A is a recent graduate from a top university with a 3.8 GPA, AI-tool certificates, and multiple AI courses. Candidate B attended an ordinary university with a 3.0 GPA but has three years of work experience: two project collapses, a data-migration incident, and countless lessons that a plan would not work.\nFive years ago, A would have had a near-certain advantage. Strong credentials, strong grades, and complete certifications were the standard signals of high capability. Today, more hiring managers hesitate—not because A is inadequate, but because credentials and certificates are losing the signals they used to convey: AI can do what they prove.\nThis is not an argument that education is useless. AI is turning textbook knowledge into a commodity, weakening credentials’ core function of screening for who has mastered that knowledge. The real differentiator is shifting toward judgment acquired through experience—tacit capabilities that cannot be written into textbooks and are difficult for AI to acquire independently.\nCredential Signals Are Failing Credentials have value not because of the paper itself, but because they play several roles: transmitting knowledge, screening ability, creating social networks, providing brand credibility, and gaining organizational recognition.\nAI primarily erodes the knowledge-transmission role. It can master what textbooks explain, often faster and more comprehensively. Its effect on ability screening and brand credibility is much more complex.\nA 2023 arXiv study by Eloundou and colleagues reported a counterintuitive result: occupations requiring more educational preparation face greater AI exposure. Workers with bachelor’s, master’s, and professional degrees have substantially higher exposure than workers without formal credentials. Programmers, tax preparers, quantitative financial analysts, and translators—traditional high-education occupations—are exactly the kinds of work LLMs are good at augmenting or replacing.\nIf AI can master the knowledge gained in a four-year degree faster and more completely, that degree’s signaling value in the labor market shrinks. Employers are shifting from “what did you study?” to “what real problem did you solve with what you learned?”\nThis is not the end of credentials, but an internal polarization. Lower-tier credentials that merely prove attendance are losing value quickly. Elite credentials are not simply appreciating; their internal structure is being reorganized:\nKnowledge transmission: AI is a substitute; codified knowledge is becoming commoditized everywhere. Learning selection: AI is an enhancer; learning ability proven through rigorous selection remains scarce. Social networks: AI cannot replace peer and alumni networks, whose value may rise. Brand credibility: prestigious institutions retain valuable screening signals amid information overload. Training environments: AI complements but cannot independently reproduce judgment forged in demanding environments. Credentials will not disappear, but their rating system is polarizing. Middle-tier credentials face the greatest devaluation pressure, while the genuinely scarce elements of elite credentials—selection, networks, environment, and brand credibility—may command higher prices.\nAI Reveals a Contradiction The mechanism behind credential devaluation is more subtle than it first appears. A February 2026 Dallas Fed study provides a useful distinction from economist Scott Davis: codified knowledge versus tacit knowledge.\nCodified knowledge is what appears in textbooks: formulas, rules, procedures, and standard operating steps. AI is highly effective at learning it. Tacit knowledge is different: you know how to act but cannot fully explain why. An experienced engineer can diagnose a fault from an engine sound; a veteran trader can sense a market turn. Such knowledge depends on long feedback cycles, specific contexts, and accumulated real outcomes.\nAI affects these forms of knowledge differently. For entry-level work dependent on codified knowledge, AI is a substitute: what newcomers spend time learning, AI can complete instantly. For experience-dependent work relying on tacit knowledge, AI is an enhancer: it handles repetitive tasks and frees experienced workers for more complex judgment.\nDavis’s data support this distinction. In occupations with the lowest experience premium, AI exposure reduced wage growth by about 0.28 percentage points. In occupations with the highest experience premium, where extensive tacit knowledge is required, exposure instead raised wage growth by about 0.2 percentage points.\nThe same technology is doing opposite things at the two ends of one skill spectrum: compressing the value of entry-level work while raising the value of senior experience.\nThe Experience Premium Is Rising Faster Credential devaluation does not mean education is unimportant. It means education’s core product—systematic codified knowledge—is being commoditized by AI. Just as steam power made physical labor cheaper, AI is making knowing things cheaper.\nExperience, by contrast, is rising in value. It is accumulated tacit knowledge: judgment that cannot be learned in class or found in textbooks, and is gained only through repeated trial and error. A doctor’s intuition after ten thousand cases, an engineer’s sense of what may fail after countless outages, or a product manager’s warning that a requirement will drift after several failed projects cannot be fully learned from public text alone.\nA 2025 Harvard study tracking employment data for 62 million U.S. workers and 285,000 firms found that entry-level positions at AI-adopting companies declined 7.7% over six quarters, while senior positions continued growing. The decline primarily reflected slower entry-level hiring rather than layoffs. The impact was U-shaped: graduates with mid-tier credentials were hurt most, while elite-university graduates and lower-educated workers were less affected. Anthropic’s March 2026 Economic Index also reported that entry rates for 22- to 25-year-olds fell about 14% in occupations with the highest AI exposure.\nThe essential issue is not simply that young people cannot find work. Experience as a screening signal is systematically becoming more important. When AI enables everyone to complete entry-level tasks, employers most want to know who has real experience.\nWhat Experience Actually Provides Experience contains four elements that AI struggles to reproduce independently. AI can learn patterns from enormous datasets, including more incident records than any individual sees. But its experience is third-person: it observes failures without bearing their consequences. Human experience is first-person: it knows not only what may fail, but what failure means. The decisive difference is simple: AI does not bear consequences.\nAn organization may deploy the most advanced AI system, but a person still signs off, takes responsibility for outcomes, and explains a failed project. The experience premium in the AI era is fundamentally a price placed on those willing to bear responsibility.\nFirst, a failure database. Experienced people know not only how to proceed, but what will not work. In complex decisions, eliminating bad options is often more important than selecting the right one. AI can learn documented failures from logs and simulations, but people encounter unfiltered failure and its subtle, difficult-to-document signals.\nSecond, contextual judgment. Textbooks describe ideal conditions. Reality has incomplete data, limited time, and conflicting stakeholders. Making decisions under ambiguity requires training in real situations. AI can increasingly analyze contextual data, but final judgment also involves value choices and responsibility.\nThird, trust networks. Knowing who to ask and which team can execute reliably is tacit organizational knowledge that remains difficult to replace.\nFourth, responsibility experience. AI can propose a plan, but cannot independently answer who will bear its consequences. A senior engineer knows the cost of failure and when experimentation is unacceptable. Only someone who has borne consequences can complete the decision chain—from “this is theoretically feasible” to “I can own the outcome if it fails.”\nThese elements cannot be fully acquired by reading; they must be accumulated through practice. As AI spreads, their scarcity becomes clearer.\nExperience Is Being Redefined Experience once roughly meant years on the job. In the AI era, that equation is loosening. Ten years spent on repetitive work that AI can also do is experience that depreciates.\nReal experience is not accumulated time, but problem density: the complexity of problems solved, the intensity of feedback received, the weight of consequences borne, and the ability to transfer lessons across contexts.\nA customer-service worker handling standard complaints for ten years may have low-value experience: low complexity and weak feedback. A product manager who fails two products in three years, adjusts a business model, confronts user churn, and makes major decisions has high-value experience because each step is an intense feedback loop.\nFuture résumés will value “years of experience” less than high-quality feedback cycles. “Led three million-user system migrations, handled two major incidents, and made key decisions across five projects” is a more persuasive signal.\nThe Winners Are AI-Enhanced Practitioners The emerging competitive profile is clear. The old path was credential → work experience. The new formula is value = foundational knowledge × AI leverage × real feedback × responsibility.\nKnowledge determines whether you understand the problem; AI determines productive efficiency; practice determines how well you understand reality; responsibility determines whether others trust you with decisions. Career growth is becoming knowledge foundation → AI augmentation → real-world closed loop.\nThe strongest people are neither pure theorists with no practice nor experienced workers who reject AI. They are AI-enhanced practitioners: people with strong fundamentals, fluent AI-tool use, and frequent opportunities for real practice. Knowledge tells them what to do, AI helps them do it better and faster, and practice accumulates judgment that textbooks lack. That judgment then improves their use of AI and acquisition of new knowledge.\nAI will not simply eliminate credentials or reward older workers. It will favor people who combine knowledge, AI, and experience to make high-quality judgments. What is being repriced is not the credential itself, but the human ability to decide in an uncertain world.\nThe Door Narrows While the Ceiling Rises Credential devaluation and experience premiums are two sides of the same coin. AI lowers the threshold for entering an industry while making sustained advancement harder. Skills that once took six months to learn can now be approached with AI in a day, but employers are less willing to pay a premium for people able to perform only entry-level work. The roles that create the most value require complex judgment and tacit experience—the capabilities newcomers lack most.\nThis creates a paradox: AI lowers barriers to entry while raising the difficulty of continuous advancement. The Industrial Revolution had a similar structure: machines shortened training, but skilled workers gained value because they knew when machines would fail. AI is recreating that pattern more quickly and at greater scale.\nA further tension is more hidden: knowledge is becoming more equal, but access to practice is becoming more unequal. A well-resourced student can start a venture, intern at a top company, and use advanced AI tools. Another may only take free courses, complete simulated projects, and never carry real responsibility. Both have ChatGPT, but they accumulate very different experience assets. AI may lower knowledge barriers while raising practice barriers.\nThink of It Like Cooking Consider two people learning to cook braised pork. One has an exceptional recipe: ingredients weighed to the gram, steps timed to the minute, and high-resolution instructions. The other follows an experienced chef, doing basic tasks and learning by watching and asking.\nIn the first week, the recipe user produces a decent dish while the apprentice is clumsy. After one month, their dishes are similar: a recipe can reliably deliver an 80-point result. A year later, the difference appears. When ingredients are not fresh or a diner requests less oil and salt, the recipe user is lost. The apprentice can judge heat from sound and timing from the color of the meat, adapting the dish anywhere from 70 to 95 points to the situation.\nThe recipe is codified knowledge. AI is more than a recipe: it resembles a super kitchen assistant with a global chef database, unlimited experimentation, and real-time feedback. Experience is the time spent beside the chef—learning how to judge heat, taste, and recover from errors. The greatest advantage after AI adoption belongs not to people who only follow recipes or reject them, but to those who use recipes while truly understanding the kitchen.\nEducation’s Rating Standard Is Changing, Not Disappearing Return to the opening hiring scenario. Candidate A’s advantage is shrinking, but credentials still matter as basic signals of learning ability, discipline, and foundational knowledge. In the AI era, however, they are increasingly an entry ticket rather than a passport: they get you through the door but do not determine how far you go.\nThe picture is not a simple one-way trend of credentials down and experience up. Credentials are polarizing: lower-tier signals are failing faster, while elite credentials retain scarce non-knowledge elements such as selection mechanisms, peer networks, and training environments. Experience is being redefined by problem density, not tenure. The biggest winners are AI-enhanced practitioners who combine foundations, AI capability, and high-frequency practice.\nThe signal determining value is shifting: from what you studied to what complex problems you solved; from test scores to judgments made under incomplete information; from what your résumé says to what irreplaceable value you can demonstrate even with AI assistance.\nThis is not the end of education. It is a rewrite of education’s underlying rating system. When knowing becomes cheap, judgment becomes expensive. When knowledge becomes a commodity, judgment earned through failure and responsibility—the capability AI currently struggles most to simulate—becomes the new scarcity. The disappearance of an old threshold does not mean everyone is equal; it means the screening standard has changed.\n","permalink":"/en/posts/ai-education-experience-premium/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\nAcademic credentials are losing their power as signals of capability because AI can master textbook knowledge. The premium is moving elsewhere: toward judgment earned through real responsibility. When knowing something is no longer scarce, what you have actually done becomes the new screening criterion.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003chr\u003e\n\u003cp\u003eTwo résumés sit before a recruiter. Candidate A is a recent graduate from a top university with a 3.8 GPA, AI-tool certificates, and multiple AI courses. Candidate B attended an ordinary university with a 3.0 GPA but has three years of work experience: two project collapses, a data-migration incident, and countless lessons that a plan would not work.\u003c/p\u003e","title":"AI and Intellectual Equality (2): Traditional Credentials Lose Value, Real Experience Gains a Premium"},{"content":" TL;DR AI is flattening basic skills: the barriers to programming, writing, and translation are approaching zero. Yet the same force is creating a deeper divide. When what you can do is no longer scarce, what you can judge becomes the true dividing line. This is not equality first and polarization later, but two sides of the same coin appearing at once.\nA junior programmer, only three months into the job, uses AI to write code and completes in one week a feature that once took a month. The code runs, the logic is coherent, and the project manager is pleased. Their output is almost indistinguishable from that of a colleague with three years of experience.\nAt the same company, a senior data analyst also uses AI. It completes data processing that used to take three days. Yet before submitting her report, she spends two extra hours checking whether the output is logically consistent, verifying the reliability of data sources, and judging whether the conclusion holds in the business context. Those two hours are precisely the hard-to-name difference between her report and a junior analyst\u0026rsquo;s.\nBoth scenes point to the same conclusion: AI lowers the threshold for basic skills while widening the gap in higher-order cognitive capabilities. What once separated people was whether they could write code. Now it is the level of problems they can solve with code.\nThis is not a temporary imbalance. It follows from AI\u0026rsquo;s technical structure.\nWhat AI Compresses Is Basic Skill Why does AI flatten basic skills before higher-order capabilities? The answer lies in how it works.\nOne useful way to understand large language models (LLMs) is as compression. They encode statistical regularities and expressive patterns from vast amounts of human knowledge into model parameters, forming a huge probabilistic prediction system. Given a request, they do not simply “create”; they produce the most likely combination based on existing patterns. This mechanism is naturally good at copying and imitation—programming conventions, writing structures, and translation patterns—which are the core of basic skills.\nWhat it does not handle as well is equally clear: creating new structures, judging subtle differences between situations, and making decisions with incomplete information. Those are the core of higher-order ability, precisely where pattern matching reaches its limits.\nSo AI\u0026rsquo;s flattening of basic skills is not accidental; it is determined by its technical structure.\nA 2025 randomized experiment by the National Bureau of Economic Research (NBER) offers more direct evidence. It divided 1,174 adults into two groups for the same business problem-solving task—one with an AI assistant and one without. AI improved everyone’s performance, but the improvement was significantly larger for people with lower educational attainment. Without AI, the gap between higher- and lower-educated participants was 0.548 standard deviations; with AI, it narrowed to 0.139—closing roughly three quarters of the initial gap. That does not mean underlying capability disappeared; AI temporarily leveled it at the execution layer.\nIn essence, AI packages the best practices of highly capable people into a tool, allowing novices to borrow that experience. It is a force of regression toward the mean: lifting the bottom rather than pushing the top higher.\nA GitHub Copilot randomized controlled trial offers evidence from another angle: developers using AI-assisted programming completed tasks more than 50% faster. With AI, a junior developer can approach the output of an intermediate developer.\nIf such augmentation can be accessed evenly, it could become the largest equalization of cognitive capability in human history. Someone who has never learned to program can use AI to write runnable code; a non-native English speaker can write a fluent English report; someone without a statistics background can complete complex data analysis. Skills that once required years of training become on-demand tools.\nThat is the reality behind the flattening of basic skills, and the optimistic vision many people hold for the AI era. Its value is helping more people cross the entry threshold. Its limit is that the world beyond that threshold is far more complex than the threshold itself.\nAfter the Threshold Flattens, Gaps Emerge from Deeper Places Between being technically available and being used evenly in practice lies a gulf.\nDigital-inequality research has a classic framework: van Dijk’s 2020 four-level model asks whether people are motivated to use technology, have access to it, have the skills to use it, and can use it effectively. The first three are questions of access; the fourth is a question of quality of use. AI is rapidly addressing the first three, while the gap in the fourth is only beginning to emerge.\nThis pattern has recurred throughout technological history. The “Year of the MOOC” in 2012 was a classic preview: platforms such as Coursera and edX promised free access to elite education for anyone worldwide, but follow-up research found that the greatest beneficiaries were still those who had already received a good education. New technology did not automatically reduce inequality; it reproduced it in a new form.\nAI is replaying the script. A 2025 PNAS study by Humlum and Vestergaard, based on large-scale Danish survey and registry data, found that ChatGPT adoption itself systematically reproduces existing inequality. Even after controlling for occupation, industry, and demographic characteristics, use gaps between groups remained. This is not theoretical speculation, but measurable reality.\nWhere do differences arise? Even if AI is free and everyone knows it exists, gaps still emerge in four dimensions: the ability to write effective prompts; the ability to assess AI output; the ability to integrate that output into a workflow; and enough domain knowledge to guide and verify AI.\nThe first layer of equalization solves the question of whether one has access. The second-layer divide comes from differences in whether people know how to use it and how effectively they use it. Surface gaps are erased; deeper ones appear.\nCognitive Offloading Is Changing How We Think There is a deeper question: even when someone has every condition needed to use AI effectively, the process itself changes how they deploy cognitive ability.\nThe key variable is cognitive offloading—handing parts of thinking over to AI. The issue is not offloading itself, but what is offloaded and how.\nA 2025 PNAS experiment by Bastani and colleagues provides causal evidence stronger than correlation. In a randomized high-school mathematics experiment, students who used AI directly to obtain answers saw their independent problem-solving ability decline. Students required to attempt reasoning first and use AI afterward retained their learning gains. The issue is not AI itself, but the usage pattern: unconstrained cognitive offloading erodes foundational ability, while structured use can protect or even strengthen learning.\nEEG research from the MIT Media Lab provides a neuroscientific clue. It found that people writing with AI showed changed patterns of brain activity: networks associated with memory and creativity were less active, while the cognitive load associated with judgment and verification increased. This is not simply a decline in ability, but a migration in how cognitive activity is distributed.\nA key distinction helps explain this shift. Earlier technologies—calculators, GPS, search engines—outsourced “nouns”: storage, retrieval, and calculation, while humans still supplied reasoning. AI outsources “verbs”: synthesis, evaluation, and judgment. When the task changes from thinking for oneself to deciding whether AI has thought correctly, the center of cognition undergoes a quiet migration. Low-value repetitive reasoning is compressed, while higher-order judgment, verification, and integration become more important.\nIts value is freeing lower-level cognitive load so people can focus on higher-level judgment. Its risk is that deciding when not to rely on AI is itself the easiest judgment to outsource. The more one relies on AI, the less one can judge when not to. This creates a cycle.\nTwo Forces Work at the Same Time Putting these clues together reveals a complete picture.\nAI applies two opposing forces to the distribution of intellectual capability.\nThe first is a compressive force: it turns programming, writing, translation, and data analysis from skills requiring years of training into on-demand tools, narrowing basic-skill gaps. The second is a stretching force: once basic skills are no longer scarce, the valuable capabilities become asking good questions, judging answer quality, integrating fragmented information, and deciding under uncertainty. These capabilities are distributed far more unevenly than basic skills.\nThe compressive force creates the appearance of equality; the stretching force creates the substance of differentiation.\nThis is the deeper meaning of the NBER experiment. Lower-educated participants greatly narrowed their gap with higher-educated participants under AI assistance, but that does not mean the latter group’s core advantage was compressed. Its value shifts away from completing routine tasks quickly and toward areas AI struggles to replace: knowing when an answer is untrustworthy, spotting implicit bias in model output, and making trade-offs in complex contexts.\nWhen the acquisition threshold for a skill is flattened, that skill loses its power to distinguish people. What once mattered was whether you could write code; now it is what level of problem you can solve with it.\nThink of It Like Driving An analogy can connect the preceding arguments.\nAI adoption resembles the adoption of cars. Cars made everyone move faster—what once took a whole day from Beijing to Tianjin now takes two hours by car. In that sense, cars achieved “mobility equality”: the speed gap from point A to point B between ordinary people and professional racers is far smaller than in the era of walking.\nBut cars also created a new gap. A racer’s value no longer lies in simply driving fast, but in controlling a vehicle under extreme conditions. Most drivers remain at the level of reaching their destination; that flattening makes the scarcity of professional drivers more visible.\nThe correspondence is clear: cars flatten the gap in whether one can get somewhere, just as AI flattens the gap in whether one can do something. The racer’s value shifts to control at the limit, just as the value of people with higher-order capability shifts to judging and verifying AI output.\nThere is a subtler layer. People who rely on cars for a long time may lose their ability to walk well. This does not mean cars are bad; it means tool substitution and capability decline are two sides of the same process. AI is similar: it narrows the gap in what most people can do, while widening the gap in the judgments they can make.\nBasic skills—coding, writing, translation, and data analysis—are driving. Higher-order skills—judgment, aesthetic sense, metacognition, and complex decision-making—are racing. AI equalizes the former while increasing the scarcity of the latter.\nDeclining walking ability is perceptible: you get out of breath after a short distance. Declining cognitive ability is not: it is hard to notice that you are thinking less. That is the warning of the EEG research—the brain changes where you cannot see it.\nAs Thresholds Disappear, Barriers Rise Return to the two opening scenes.\nThe junior programmer finishes a feature with AI, but their ceiling remains at being able to write code with AI. The senior analyst uses AI to process data, but her judgment, domain intuition, and critical validation are what actually deliver value.\nAI adoption is not a simple act of intellectual equalization. It silently rewrites the standards of selection. The old distinction—what you know how to do—is losing force. New distinctions are forming: what questions you can ask, when you can refuse to trust AI, and what kind of judgment you can form from fragments.\nThese abilities are much more unevenly distributed than programming and writing skills. AI’s real impact on intellectual gaps is neither simple narrowing nor widening, but restructuring: it erases one old layer of inequality and opens a new one at a deeper level.\n","permalink":"/en/posts/ai-intelligence-equality/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\nAI is flattening basic skills: the barriers to programming, writing, and translation are approaching zero. Yet the same force is creating a deeper divide. When \u003cem\u003ewhat you can do\u003c/em\u003e is no longer scarce, \u003cem\u003ewhat you can judge\u003c/em\u003e becomes the true dividing line. This is not equality first and polarization later, but two sides of the same coin appearing at once.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003chr\u003e\n\u003cp\u003eA junior programmer, only three months into the job, uses AI to write code and completes in one week a feature that once took a month. The code runs, the logic is coherent, and the project manager is pleased. Their output is almost indistinguishable from that of a colleague with three years of experience.\u003c/p\u003e","title":"AI and Intellectual Equality (1): Basic Skills Flatten, Capability Gaps Widen"},{"content":" TL;DR: The three paths ultimately converge because real tasks simultaneously require understanding the current situation, reusing historical experience, and executing real actions. The endgame for agents is not being better at conversation, but becoming a governed goal-completion system: understanding context, consolidating experience, calling tools, executing tasks, and accepting permissions, auditing, rollback, and user control.\nThis is the final article in the Agent Evolution Series.\nIn the previous four articles, we\u0026rsquo;ve already deconstructed the three paths separately: the execution path, the self-evolution path, and the personal context path, as well as the final form they may form after convergence. The first three articles were more like looking down each path: what it is, why it emerged, what stages it goes through, and where it ultimately leads. The fourth article put the three paths into the same system, discussing how they combine into a governed Agent Runtime.\nNow in this article, we need to raise the perspective one more level.\nThis article will no longer just repeat the introduction of the three paths, but answer a more global question: Why do agents simultaneously develop along these three paths? Why do they look like different product directions in the short term, but inevitably converge in the long term? Why is the endgame for agents not a better chatbot, but a governed goal-completion system?\nIf the first four articles are local deconstruction, then this one is the global synthesis.\n1. Putting the Three Paths Back on the Same Map The evolution of agents is essentially the process of AI moving from \u0026ldquo;answering questions\u0026rdquo; to \u0026ldquo;completing goals.\u0026rdquo;\nOnce viewed from this perspective, the three paths are not isolated product categories, but three indispensable capabilities within the same goal-completion system.\nThe first is the Execution Path. It answers: Can the agent actually do things? Past AI assistants mainly gave advice: telling users how to fix code, write emails, handle errors, organize materials. But the execution path goes a step further, connecting agents to tools, browsers, file systems, shells, APIs, and business systems to directly complete tasks within user authorization. This path starts with tool calling, moves into browser and computer use, then local automation, and finally becomes the action layer in a mature agent system. Its value is most direct because as soon as an agent completes a real action, the user immediately feels time saved. But its risk is also most direct: once an agent can operate real systems, errors are no longer just wrong answers—they can become misoperations, unauthorized access, data leaks, or irreversible losses. So the endpoint of the execution path is not \u0026ldquo;more tools are better,\u0026rdquo; but \u0026ldquo;reliably completing tasks within strict boundaries.\u0026rdquo;\nThe second is the Self-Evolution Path. It answers: Can the agent get better with use? The problem with many AI assistants is not weak single-use capability, but that every time is like the first cooperation. The user explains project background, tool paths, code conventions, and workflows once, and has to re-explain them next time. The self-evolution path addresses this lack of continuity. It starts with in-conversation learning, moves to long-term memory, then skills, and finally enters evaluation, version management, and rollback stages. Truly valuable self-evolution is not about mysterious self-awakening, but about consolidating effective experience from historical tasks into reusable, checkable, deletable, rollbackable capability assets. The risk of this path is also clear: if the agent learns wrong, the error is no longer a one-time mistake but long-term contamination. Wrong memories, wrong skills, outdated processes, and wrong task trajectories will repeatedly affect judgment in future tasks. So the endpoint of the self-evolution path is not \u0026ldquo;automatically remember everything,\u0026rdquo; but \u0026ldquo;evaluably, versionably, governably consolidate capability.\u0026rdquo;\nThe third is the Personal Context Path. It answers: Can the agent truly understand the user and organization? A general model can be very smart, but by default it doesn\u0026rsquo;t know who the user is, what project they\u0026rsquo;re working on, who they collaborate with, which approaches have been rejected, which tasks are most urgent, or which information is sensitive. The personal context path addresses the gap between general intelligence and specific situations. It starts with simple preference memory, moves to workspace context, then connects email, calendar, documents, code repos, chat history, task systems, and other multi-source data, ultimately forming an explainable, editable, deletable, migratable personal memory system. This path has high long-term value because when model capabilities become increasingly commoditized, what\u0026rsquo;s truly hard to replicate is the user\u0026rsquo;s specific context. Models can be called by multiple products, but the user\u0026rsquo;s projects, relationships, historical decisions, work methods, and long-term preferences cannot be easily copied. But its prerequisite is trust. Users won\u0026rsquo;t entrust their email, documents, calendar, chat, code, and task systems to an unexplainable, uncontrollable, undeletable black box. So the endpoint of the personal context path is not \u0026ldquo;connect more data sources,\u0026rdquo; but \u0026ldquo;build a trustworthy understanding layer.\u0026rdquo;\nPutting these three paths together, they correspond to three layers of a complete agent system:\nThe personal context path provides the understanding layer; The self-evolution path provides the learning layer; The execution path provides the action layer. The understanding layer lets the agent know the current situation, the learning layer lets the agent reuse past experience, and the action layer lets the agent turn goals into real results.\nThis is the structure that the first four articles collectively point to.\n2. The Final Form Is Not a Chatbot, but Agent Runtime After the three paths converge, the final form of agents will not be a chatbot that\u0026rsquo;s better at conversation.\nIt\u0026rsquo;s more like an Agent Runtime.\nThe word Runtime is important. It means the agent is not a single model or single interface, but a runtime system organized around goal completion.\nIn this system, the model is just one core component. What truly makes the agent usable is the entire structure around the model: context, memory, skills, task planning, tool execution, permission control, audit logs, rollback mechanisms, and user feedback.\nA mature Agent Runtime roughly contains eight layers of capability:\nContext layer: Connects personal and organizational data, judges which information is relevant to the current task; Memory layer: Saves long-term stable information, while allowing users to view, modify, and delete; Skills layer: consolidates repetitive workflows into reusable skills; Planning layer: Breaks user goals into executable steps; Action layer: Calls browsers, files, shells, APIs, and business systems to complete operations; Permission layer: Security boundary determining what the agent can do; Audit and rollback layer: Ensures operations are traceable and recoverable; Feedback layer: Sends task results back to improve future performance. So the core of Agent Runtime is not \u0026ldquo;being more human-like,\u0026rdquo; but \u0026ldquo;being a reliable goal-completion system.\u0026rdquo;\nIt may not replace all software. The more likely form is that agents become a coordination layer above software.\nToday, users manually switch between different software: receiving requirements in email, checking background in documents, looking at time in calendar, communicating in Slack or Teams, modifying code in GitHub, updating tasks in Linear or Jira, searching in browsers, running commands in local terminals.\nAgent Runtime\u0026rsquo;s value is coordinating these tools around user goals.\nUsers no longer start with \u0026ldquo;which software should I open,\u0026rdquo; but with \u0026ldquo;what goal do I want to accomplish.\u0026rdquo; The agent maps goals to context, memory, skills, tools, and permissions, advancing tasks within controllable boundaries.\nThis is the form the three paths ultimately converge into.\n3. Why Agents Won\u0026rsquo;t Stop at Chat Assistants To understand why agents continue to evolve, we must first see a fact: what users truly need is not chat, but goal completion.\nUsers ask AI to write emails not to get text, but to complete communication. Users ask AI to analyze errors not to get an explanation, but to fix problems. Users ask AI to summarize materials not for the summary itself, but to make judgments, write articles, make plans, or advance projects.\nChat is just the entry point. In the early stage, the chat entry point is important enough because it allows people to call model capabilities with natural language. But when models can stably generate high-quality text, user needs naturally push forward.\nUsers will continue to ask: Since you know how, can you just do it for me? Can you remember this experience next time? Can you understand my project background instead of making me explain it every time? Can you enter my real workflow and continuously help me advance tasks?\nThese questions together push agents beyond the boundaries of chat assistants.\nChat solves the expression problem; agents must solve the goal completion problem.\nThis is also why agent evolution necessarily involves tools, memory, context, permissions, and governance. Just a bigger input box cannot accomplish this leap.\n4. Why the Paths Diversified Agents diversified into three paths not because the industry likes creating concepts, but because real tasks can themselves be broken into three questions.\nThe first question: Tasks need to be executed. If agents can\u0026rsquo;t act, no matter how well they answer, the execution cost remains with the user. Users still need to copy, paste, open web pages, run commands, modify files, and submit tasks themselves. This drove the execution path.\nThe second question: Experience needs to be consolidated. If agents can\u0026rsquo;t learn from historical tasks, they remain beginners every time. Project backgrounds, failure experiences, tool conventions, and workflows that users explained once can\u0026rsquo;t be converted into advantages for the next collaboration. This drove the self-evolution path.\nThe third question: Action needs context understanding. If agents don\u0026rsquo;t understand user and organizational context, even if they can act, they may go in the wrong direction. They don\u0026rsquo;t know which information is important, which actions are sensitive, which approaches have been rejected, which preferences are long-term stable. This drove the personal context path.\nSo the three paths are not forcibly extracted from technical terminology, but naturally grown from the structure of real tasks.\nComplex tasks always require three things: understanding the current situation, leveraging past experience, and executing real actions.\nUnderstanding, learning, and acting—these three things combined make a complete agent.\n5. Why Short-Term Diversification If the three paths ultimately converge, why did they diverge in the first place?\nThe reason is that the entry paths and maturity conditions of the three paths are different.\nThe execution path most easily demonstrates short-term value. As soon as an agent successfully completes one file organization, web search, code modification, or form fill, the user immediately feels the value. The execution path is naturally suited for demos and for starting with low-risk, clearly bounded tasks.\nThe self-evolution path is different. It\u0026rsquo;s hard to prove itself with one demo. Whether an agent truly gets better with use requires long-term observation: does it reduce user repetition, avoid repeated errors, improve success rates for similar tasks, and consolidate historical experience into reusable capabilities? So this path is more infrastructure-oriented and more dependent on evaluation, version management, and rollback mechanisms.\nThe personal context path has yet another rhythm. Its value is high, but requires trust as a prerequisite. Users won\u0026rsquo;t initially give an agent access to all their email, documents, calendar, chat, code repos, and task systems. It must first prove itself explainable, editable, deletable, and migratable—preferably local-first with permission layering—before users gradually open more context.\nTherefore, short-term diversification is not because the three paths are mutually exclusive, but because they suit different entry points.\nExecution first proves efficiency, self-evolution consolidates capability, personal context builds trust.\nThis is the real reason for short-term diversification.\n6. Why Long-Term Convergence Is Inevitable Short-term diversification is possible, but long-term independence is hard.\nThe reason is also simple: any single path will hit a ceiling.\nAn agent with only execution lacks judgment basis. It may be able to call tools, operate web pages, modify files, run commands, but without understanding user background or consolidating historical experience, it easily becomes a high-risk automation script. It can do things, but doesn\u0026rsquo;t necessarily know what\u0026rsquo;s worth doing or how to align with the user\u0026rsquo;s situation.\nAn agent with only self-evolution lacks verification scenarios. It may be able to write memories, generate skills, summarize task trajectories, but without a real execution loop, it\u0026rsquo;s hard to judge whether these experiences are actually effective. Without context control, it may also apply correct experiences to wrong scenarios.\nAn agent with only personal context lacks an action loop. It may understand the user well, knowing project background, historical decisions, and task status, but if it can\u0026rsquo;t execute or consolidate repetitive workflows into skills, it stays at the knowledge base, search, or contextual Q\u0026amp;A tool stage. It can understand you, but can\u0026rsquo;t truly help advance tasks.\nSo in the long term, the three paths must converge.\nThe understanding layer provides background, the learning layer provides experience, the action layer completes operations. Combined, agents can go from \u0026ldquo;can talk\u0026rdquo; to \u0026ldquo;can do,\u0026rdquo; from \u0026ldquo;one-shot assistant\u0026rdquo; to \u0026ldquo;long-term collaborator,\u0026rdquo; and from \u0026ldquo;general model\u0026rdquo; to \u0026ldquo;user-understanding productivity system.\u0026rdquo;\n7. Why the Final Form Must Be Governed The stronger the agent\u0026rsquo;s capabilities, the more important governance becomes.\nThis is the most easily underestimated but most critical point in the entire agent evolution.\nA model that only chats mainly risks wrong answers. Wrong answers can be ignored, corrected, or re-asked.\nBut a mature agent has many more capabilities: reading context, saving memories, generating skills, calling tools, operating files, running commands, accessing business systems, and adjusting future behavior based on historical experience.\nThese capabilities upgrade risks from \u0026ldquo;saying wrong things\u0026rdquo; to \u0026ldquo;doing wrong things.\u0026rdquo;\nMisoperations, unauthorized access, data leaks, prompt injection, wrong memories, wrong skills, privacy exposure, compliance issues, and unclear responsibility attribution all become real risks.\nTherefore, the final form of agents cannot be an infinitely free autonomous entity.\nIt must be a governed system.\nGovernance is not an external restriction imposed on agents, but the prerequisite for agents to enter real productivity scenarios. Without least privilege, enterprises won\u0026rsquo;t trust agents to access business systems; without human confirmation, users won\u0026rsquo;t trust agents to execute high-risk actions; without audit logs, there\u0026rsquo;s no way to know what the agent did; without rollback mechanisms, erroneous operations can\u0026rsquo;t be remedied; without memory editing and skill version management, self-evolution can become self-contamination; without data access control, personal context becomes a privacy risk.\nSo the keywords for mature agents are not \u0026ldquo;fully autonomous,\u0026rdquo; but \u0026ldquo;authorizable, auditable, rollbackable, controllable.\u0026rdquo;\nThis is also why the final form is a governed Agent Runtime, not a boundaryless AI assistant.\n8. Why Model Capability Is Not the Only Moat Many people simplify the future of agents to models getting stronger.\nModels are certainly important. Without strong enough models, agents can\u0026rsquo;t understand tasks, plan steps, call tools, handle exceptions.\nBut models aren\u0026rsquo;t everything.\nA truly usable agent also depends on high-quality personal and organizational context, reusable skills, stable tool connections, permission governance, audit rollback, memory control, evaluation loop, and user trust.\nModels are more like an engine. But a roadworthy system can\u0026rsquo;t just have an engine—it also needs a chassis, brakes, navigation, dashboard, safety systems, and maintenance mechanisms.\nMany agent demos look impressive but are hard to actually deploy for this reason: they showcase the model\u0026rsquo;s instantaneous capability but don\u0026rsquo;t solve long-term system capabilities.\nFrom this perspective, the future moat for agents is not just \u0026ldquo;who uses a stronger model,\u0026rdquo; but: Who has higher quality context, who can consolidate experience into governable skills, who can stably execute in complex tool environments, and who makes permissions, auditing, rollback, and user control into foundational capabilities.\nThis is the key to agents moving from proof-of-concept to productivity infrastructure.\n9. From a Global View, Agents Are a Coordination Layer Above Software If you only look at individual products, agents are easily understood as some kind of new application.\nBut from a longer cycle, agents are more likely to become a coordination layer above software.\nToday\u0026rsquo;s software world is very fragmented. A real task often spans multiple systems: requirements in email, background in documents, communication in chat tools, time in calendar, code in GitHub, tasks in Linear or Jira, execution in local terminals, supplementary information in browsers.\nIn the past, users manually switched between these systems, manually transferred information, manually judged next steps, manually executed actions.\nThe long-term value of agents is transforming this cross-software coordination into a goal-centric workflow.\nUsers propose goals, agents understand context, call historical experience, select appropriate tools, execute actions within permission boundaries, and feed results back to users.\nIt\u0026rsquo;s not about eliminating all software, but making collaboration between software more natural.\nThis is also why Agent Runtime is more accurate than \u0026ldquo;super chatbot.\u0026rdquo; A super chatbot is still an entry point; Agent Runtime is a layer of cross-software, cross-data, cross-task coordination infrastructure.\n10. Why Commercialization Starts with Small Scenarios Although the final form is large, agent commercialization won\u0026rsquo;t start with the most complex, highest-risk tasks.\nIt will definitely start with clearly bounded, risk-controllable, value-measurable small scenarios.\nThis is determined by three real constraints.\nFirst, users need to see ROI. If an agent can\u0026rsquo;t clearly save time, reduce costs, or improve quality, it easily stays at proof-of-concept.\nSecond, enterprises need to control risk. High-permission, high-risk, irreversible tasks won\u0026rsquo;t initially be handed to agents. Enterprises are more likely to start with ticket processing, internal knowledge retrieval, code assistance, low-risk process automation.\nThird, personal context needs trust accumulation. Users won\u0026rsquo;t immediately open all data, but start with partial authorization. Let the agent handle a certain project, certain type of document, certain workspace, confirm it\u0026rsquo;s controllable, deletable, explainable, then gradually expand the scope.\nSo in the short term, we\u0026rsquo;ll see many agents entering specific scenarios: file organization, web search, email drafts, meeting notes, code assistance, ticket processing, knowledge base Q\u0026amp;A, personal knowledge management, team memory.\nThese may not look like the final form, but they\u0026rsquo;re the entry points to the final form.\nAs reliability, permissions, auditing, context, and skill governance gradually mature, agents will enter more core business processes.\n11. Conclusion: Three Paths, Three Faces of One System Returning to the core question of the entire series: Where are agents heading?\nThe answer is not a single path winning out.\nThe execution path, self-evolution path, and personal context path look like three directions in the short term, but in the long term, they\u0026rsquo;re three faces of one system.\nThe execution path solves the action problem: how agents turn goals into real actions. The self-evolution path solves the learning problem: how agents consolidate capability from historical tasks. The personal context path solves the understanding problem: how agents understand user and organizational situations.\nThe final form solves the governance problem: how these capabilities combine under permissions, auditing, rollback, and user control into a trustworthy system.\nSo, the endgame for agents is not being better at conversation, but becoming a governed goal-completion system.\nIt understands context, consolidating experience, calls tools, executes tasks, and throughout the process accepts permissions, auditing, rollback, and user control.\nThis is also the fundamental logic of agents evolving from chat assistants to productivity infrastructure.\n","permalink":"/en/posts/agent-evolution-convergence/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: The three paths ultimately converge because real tasks simultaneously require understanding the current situation, reusing historical experience, and executing real actions. The endgame for agents is not being better at conversation, but becoming a governed goal-completion system: understanding context, consolidating experience, calling tools, executing tasks, and accepting permissions, auditing, rollback, and user control.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis is the final article in the Agent Evolution Series.\u003c/p\u003e\n\u003cp\u003eIn the previous four articles, we\u0026rsquo;ve already deconstructed the three paths separately: the execution path, the self-evolution path, and the personal context path, as well as the final form they may form after convergence. The first three articles were more like looking down each path: what it is, why it emerged, what stages it goes through, and where it ultimately leads. The fourth article put the three paths into the same system, discussing how they combine into a governed Agent Runtime.\u003c/p\u003e","title":"Agent Evolution Series (5): A Global View — Why the Three Paths Ultimately Converge"},{"content":" TL;DR: The final form of agents is not one path winning over the others, but the convergence of execution, self-evolution, and personal context into a governed Agent Runtime. This system consists of an understanding layer, a learning layer, an action layer, and a governance layer. Key capabilities include context, memory, skills, tool execution, permissions, auditing, and rollback.\nThis is the fourth article in the Agent Evolution Series.\nThe first three articles discussed three paths:\nExecution path: agents moving from answering to doing; Self-evolution path: agents moving from one-shot assistants to long-term collaborators; Personal context path: agents moving from general assistants to assistants that understand users. This article discusses the final form after the three paths converge.\nThe conclusion is:\nThe final form of agents won\u0026rsquo;t be a single path winning out, but the three paths converging into a governed Agent Runtime.\nThe focus of this final form is not explaining \u0026ldquo;why this happens,\u0026rdquo; but illustrating \u0026ldquo;what it will look like and how each layer develops.\u0026rdquo;\n\u0026ldquo;Why this happens\u0026rdquo; is reserved for the fifth article.\n1. Why the Final Form Is Not a Single Point Capability Execution, self-evolution, and personal context each solve different problems.\nExecution solves: Can the agent do things? Self-evolution solves: Can the agent get better with use? Personal context solves: Can the agent understand the user?\nBut any single path is incomplete.\n1. Only Execution, Not Enough An agent that only executes can call tools, modify files, run commands, and operate browsers. But if it doesn\u0026rsquo;t understand user context or consolidate long-term experience, it becomes a high-risk automation script. It can do things, but doesn\u0026rsquo;t necessarily know whether it should, why, or how to align with user habits.\n2. Only Self-Evolution, Not Enough An agent that only emphasizes learning can write memories, generate skills, and summarize task trajectories. But without clear execution scenarios, its learning is hard to verify. Without governance mechanisms, it may also permanently entrench erroneous experience. It can accumulate, but not necessarily correctly.\n3. Only Personal Context, Not Enough An agent that only understands the user can comprehend email, documents, calendar, tasks, and personal preferences. But if it can\u0026rsquo;t execute or consolidate experience into skills, it stays at the personal knowledge base or contextual Q\u0026amp;A tool stage. It understands you, but can\u0026rsquo;t necessarily help you get things done.\n2. The Final Form: Agent Runtime A mature agent is more like a runtime system than a single chat window.\nIt needs to simultaneously handle: user goals, context, memory, skills, planning, tools, permissions, auditing, rollback, and feedback.\nThis is the Agent Runtime.\nIt\u0026rsquo;s not a model, but a system. The underlying model is just one component.\n3. Core Layers of Agent Runtime A mature Agent Runtime will contain at least eight layers.\n1. Context Layer: Understanding the User and Organizational Environment The context layer comes from the personal context path. It connects and organizes user or organizational data: email, calendar, documents, code repos, chat history, task systems, local files, cloud storage, enterprise internal systems. The context layer doesn\u0026rsquo;t just sync data; it judges which data is relevant to the current task, which data can be used, which is sensitive, and which is outdated. Its goal is to prevent agents from understanding users from zero each time.\n2. Memory Layer: Saving Long-Term Useful Information The memory layer saves stable information: user preferences, project background, team conventions, historical decisions, common errors, verified workflows. A mature memory layer must be explainable, editable, deletable, and migratable. If memory is uncontrollable, users won\u0026rsquo;t trust it.\n3. Skills Layer: Turning Experience into Reusable Capabilities The skills layer comes from the self-evolution path. It consolidates historical experience into skills. Skills can include: work steps, tool descriptions, applicable conditions, examples, checklists, risk warnings, failure handling methods, and evaluation criteria. A mature skills layer must have lifecycle management: creation, testing, versioning, rollback, deletion, and auditing.\n4. Planning Layer: Breaking Goals into Executable Steps The planning layer converts user goals into task plans. It needs to judge: what context the current task needs, which memories to load, whether a skill is needed, which tools to call, which steps have higher risk, which actions require human confirmation, and how to retry or reroute after failure. The planning layer is also where multi-model scheduling may occur. Different models can handle planning, retrieval, code generation, vision understanding, summarization, and evaluation.\n5. Action Layer: Calling Tools to Complete Real Operations The action layer comes from the execution path. It calls tools and executes real operations: file system, shell, browser, API, database, SaaS systems, enterprise internal business systems. The action layer is where users most easily perceive value. But it must also be the most strictly restricted. Because from this layer onward, model outputs become real consequences.\n6. Permission Layer: Determining What the Agent Can Do The permission layer is the security boundary of Agent Runtime. It determines: what data the current task can access, whether it can write files, execute commands, send external messages, access production systems, which operations require user confirmation, and whether it complies with enterprise policy. A mature agent\u0026rsquo;s permission system will increasingly resemble a combination of OS permissions, enterprise IAM, and automated approval workflows.\n7. Audit and Rollback Layer: Making Operations Traceable and Recoverable Once agents enter real workflows, they must leave traceable records. The audit layer needs to record: user goals, context used, tools called, operations executed, files modified, data accessed, user confirmation records, failure and retry processes. The rollback layer is responsible for undoing erroneous operations as much as possible. Without audit and rollback, agents will find it hard to enter enterprise or high-risk personal scenarios.\n8. Feedback Layer: Enabling Continuous Improvement The feedback layer sends task results and user corrections back to the memory, skills, and planning layers. But feedback can\u0026rsquo;t be blindly written into long-term memory. The system must judge: is this experience reusable? Is it a one-time occurrence? Does it conflict with old memories? Does it need user confirmation? Should a new skill be generated? Should an old skill be modified? The feedback layer is where self-evolution truly happens.\n4. Development Path of Agent Runtime The final form won\u0026rsquo;t appear suddenly, but will develop gradually.\n1. From Single Tools to a Tool Ecosystem Early agents will only call a few tools. Later, unified tool protocols, tool markets, tool permissions, and tool auditing will emerge. Protocols like MCP represent the direction of standardizing connections between models, tools, and data sources.\n2. From Temporary Context to Long-Term Context Early agents mainly rely on current conversation context. Gradually, project context, personal context, organizational context, and long-term memory will be introduced. Context will evolve from \u0026ldquo;temporary input\u0026rdquo; to \u0026ldquo;long-term asset.\u0026rdquo;\n3. From Manual Processes to Skills Many repetitive tasks initially rely on manual explanation by users. Later, these processes will be consolidated into skills. Skills will become the reusable capability units of Agent Runtime.\n4. From Free Execution to Governed Execution Early demos often emphasize how much agents can do. Mature systems emphasize what conditions agents cannot do things under. Permissions, approval, auditing, sandboxes, and rollback will become the infrastructure of execution.\n5. From Single Assistant to Cross-Software Coordination Layer Ultimately, agents may not replace all software. They\u0026rsquo;re more likely to become a coordination layer above software. Users will no longer start with \u0026ldquo;which software should I open,\u0026rdquo; but with \u0026ldquo;what goal do I want to accomplish.\u0026rdquo; Agent Runtime maps goals to context, skills, tools, permissions, and operations.\n5. Differences Between Personal and Enterprise Agent Final Forms Agent Runtime will have different emphases in personal and enterprise scenarios.\nPersonal Agent Personal agents emphasize: personal context, local-first, data ownership, coordination of schedule, email, documents, code, and tasks, and viewable, editable, deletable, migratable memory. The core value of personal agents is reducing repetitive explanation and improving personal workflow efficiency.\nEnterprise Agent Enterprise agents emphasize: organizational knowledge, role-based permissions, audit compliance, business processes, security policies, ROI measurement, and responsibility attribution. Enterprises won\u0026rsquo;t accept unauditable, uncontrollable, non-rollbackable agents. The core of enterprise agents is not making AI freer, but improving process efficiency within clear boundaries.\n6. Judging Whether an Agent Is Close to the Final Form Judging whether an agent is close to the final form shouldn\u0026rsquo;t just look at how strong the model is or how many tools it has.\nMore importantly, does it have:\nAbility to understand personal or organizational context; Ability to preserve long-term memory with user control; Ability to consolidate experience into skills; Ability to decompose tasks and select appropriate tools; Ability to execute real operations within permission boundaries; Ability to distinguish low-risk and high-risk actions; Ability to audit, explain, and roll back critical operations; Ability to improve from feedback without self-contamination; Ability to connect different software and data sources while maintaining least privilege; Ability to build trust in personal and enterprise scenarios. A system with these capabilities is no longer just a chatbot, but an Agent Runtime.\n7. Summary The final form of agents is not a single path winning out, but the convergence of three paths.\nThe execution path provides the action layer. The self-evolution path provides the learning layer. The personal context path provides the understanding layer.\nTogether, they form a governed Agent Runtime.\nIt can understand users, consolidate experience, call tools, respect boundaries, improve efficiency, and let users know what it did, why, and whether it can be undone.\nThe next article will no longer repeat the introduction of the three paths, but answer a more fundamental question: why agents necessarily develop along these three paths and ultimately converge.\n","permalink":"/en/posts/agent-evolution-final-form/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: The final form of agents is not one path winning over the others, but the convergence of execution, self-evolution, and personal context into a governed Agent Runtime. This system consists of an understanding layer, a learning layer, an action layer, and a governance layer. Key capabilities include context, memory, skills, tool execution, permissions, auditing, and rollback.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis is the fourth article in the Agent Evolution Series.\u003c/p\u003e\n\u003cp\u003eThe first three articles discussed three paths:\u003c/p\u003e","title":"Agent Evolution Series (4): The Final Form — Governed Agent Runtime"},{"content":" TL;DR: The core of the personal context path is transforming agents from general assistants into assistants that truly understand your situation. It will evolve from preference memory to workspace context, multi-source data integration, and controllable personal memory systems. The long-term moat isn\u0026rsquo;t connector count, but trustworthy, explainable, editable, deletable context capability.\nThis is the third article in the Agent Evolution Series.\nThe first discussed the execution path: how agents move from answering to doing. The second discussed the self-evolution path: how agents move from one-shot assistants to long-term collaborators.\nThis article discusses the third path: The Personal Context Path.\nThis path asks:\nCan an agent truly understand who the user is, what they\u0026rsquo;re working on, what relationships and task networks they\u0026rsquo;re in, and what information is truly important for the current task?\nIf the execution path solves \u0026ldquo;doing\u0026rdquo; and the self-evolution path solves \u0026ldquo;growth,\u0026rdquo; the personal context path solves \u0026ldquo;understanding.\u0026rdquo;\nProjects like OpenHuman represent this path. They emphasize local-first, personal data syncing, SQLite, Markdown, Memory Tree, Obsidian-style knowledge bases, and integration of Gmail, Notion, GitHub, Slack, Calendar, Drive, Linear, Jira, and other data sources.\n1. What Is the Personal Context Path The core goal of the personal context path is making agents go from general assistants to \u0026ldquo;assistants that know me.\u0026rdquo;\nA general model can answer many questions, but by default it doesn\u0026rsquo;t know the user.\nIt doesn\u0026rsquo;t know:\nWhat project the user is currently working on; Who they\u0026rsquo;ve recently discussed things with; Which tasks are completed; Which approaches have been rejected; What expression style the user prefers; How the team collaborates internally; What conventions a certain repository has; Which information is sensitive; What background is relevant to the current task. So users often need to repeatedly provide context.\nThe personal context path aims to solve this repetitive explanation.\nIt wants agents to understand the user\u0026rsquo;s long-accumulated digital environment, including email, documents, calendar, chat, code, task systems, and personal preferences.\n2. Why the Personal Context Path Emerged The fundamental reason is that general intelligence doesn\u0026rsquo;t equal specific understanding.\nA model can be very powerful, but by default it doesn\u0026rsquo;t know who the user is or what task, relationship, and information environment the user is currently in.\nIt might answer a generally correct question, but the answer may not suit the current user.\nSo as model capabilities grow stronger, users increasingly realize another problem:\nYour answers are good, but you don\u0026rsquo;t understand my specific situation.\nThis is why the personal context path emerged.\nWhat users truly need is not an abstractly smart model, but an assistant that understands their situation.\nFor example:\nJudging task urgency based on the user\u0026rsquo;s calendar; Understanding the context of a requirement from email threads; Determining development task background from code repos and issues; Knowing which approaches have already been discussed from historical documents; Generating content closer to the user\u0026rsquo;s voice; Determining who needs to be informed based on team collaboration relationships. This path develops along \u0026ldquo;preference memory → workspace context → multi-source personal data integration → controllable personal memory system\u0026rdquo; for internal reasons.\nBecause for an agent to truly understand the user, it must first remember stable preferences, then understand the workspace, then connect more complete data sources, and finally solve trust and control issues.\nSo the long-term value of the personal context path is not making agents better at answering general questions, but making their judgments and actions more tailored to specific users.\nAs underlying model capabilities become increasingly commoditized, personal and organizational context will become harder-to-migrate, more differentiating assets.\n3. Stages of the Personal Context Path The personal context path roughly goes through four stages.\n1. Preference Memory The earliest personal context is usually simple preference memory.\nFor example:\nThe user prefers Chinese responses; The user prefers concise and direct answers; The user prefers a certain code style; The user frequently does certain types of tasks. This type of memory can improve the experience, but it\u0026rsquo;s only the shallowest layer of personal context.\nIt solves expression and interaction preferences, not enough to support complex task understanding.\n2. Workspace Context The second stage is workspace context.\nAgents begin to understand the user\u0026rsquo;s current project, documents, code repos, task systems, and team materials.\nAt this point, they no longer just know user preferences—they also know what work the user is handling.\nFor example:\nCurrent project goals; Relevant documents; Issue status; PR discussions; Team conventions; Meeting notes; Task priorities. This stage significantly improves agent relevance.\n3. Multi-Source Personal Data Integration The third stage is multi-source personal data integration.\nAgents connect more data sources: Gmail, Calendar, Slack, Notion, GitHub, Drive, Linear, Jira, local files, browser data.\nConnecting more data sources improves context completeness, but also brings privacy, security, and permission issues.\nThe core challenge of this stage is no longer \u0026ldquo;can we connect,\u0026rdquo; but \u0026ldquo;can we use it correctly.\u0026rdquo;\n4. Controllable Personal Memory System The fourth stage is a controllable personal memory system.\nA mature personal context system can\u0026rsquo;t just collect data—it must let users control their data.\nIt should support:\nViewing what the system has remembered; Modifying incorrect memories; Deleting memories that shouldn\u0026rsquo;t be kept; Exporting and migrating context; Restricting different tasks to access different data; Managing sensitive information in layers. The goal of this stage is building trust. Without trust, the more complete the personal context, the more uneasy the user.\n4. Personal Context Is Not a Connector Competition The personal context path is easily misunderstood as \u0026ldquo;connecting more apps.\u0026rdquo;\nSo products emphasize how many tools they connect: email, calendar, documents, cloud storage, chat, code repos, task systems.\nBut connection is just the first step.\nWhat\u0026rsquo;s truly hard is turning this data into trustworthy, usable, controllable context.\nKey questions include:\nWhich information is worth keeping long-term? Which information is only suitable for the current task? Which information is outdated? Which information conflicts? Which information is sensitive? What context does the current task actually need? Why did the agent reference this memory? Can the user correct errors? So the competitive core of the personal context path is not connector count, but context governance capability.\n5. Why Local-First Matters The more complete the personal context, the more sensitive the data.\nIf an agent can access email, calendar, chat history, documents, code repos, cloud storage, and task systems, it may possess a very complete picture of the user\u0026rsquo;s digital life.\nAt this point, users naturally ask:\nWhere is my data stored? What content is synced? What content is summarized into memories? Can the model see all data? Is deletion thorough? Can I export? Can I migrate? Can different tasks only use necessary context? This is where the local-first approach matters.\nLocal-first doesn\u0026rsquo;t mean all computation must be offline, nor does it prohibit using cloud models.\nWhat it emphasizes is user control over data.\nSQLite, Markdown, Memory Tree, Obsidian-style knowledge bases—these designs matter because they turn personal context from an invisible black box into assets users can view, edit, and migrate.\n6. Trustworthy Memory Is the Core of the Personal Context Path The most critical capability of the personal context path is trustworthy memory.\nTrustworthy memory includes at least four aspects.\n1. Explainable Users should know what the agent has remembered. When an answer is influenced by a memory, the system should ideally explain where that memory came from and why it\u0026rsquo;s relevant.\n2. Editable Memories can be wrong. If the agent misunderstands the user, the user should be able to directly correct it, rather than letting the error affect future tasks long-term.\n3. Deletable Some information shouldn\u0026rsquo;t be kept long-term. Users need a clear deletion entry point, and deletion should truly take effect.\n4. Migratable Personal context is the user\u0026rsquo;s asset and shouldn\u0026rsquo;t be locked into a platform. If you can\u0026rsquo;t export or migrate, long-term context becomes a product lock-in mechanism.\n7. Permission Layering for Personal Context Not all tasks need all context.\nFor example:\nWriting a public article doesn\u0026rsquo;t need to read private email; Modifying code doesn\u0026rsquo;t necessarily need calendar access; Scheduling a meeting needs contacts and schedule, but not private repos; Summarizing project progress may need issues, docs, and meeting notes, but not private chat history. A mature personal context system must support minimum necessary context.\nThis means:\nDifferent tasks access different data; Different models get different context; Sensitive data is isolated by default; High-risk data access requires confirmation; Enterprise scenarios must also comply with organizational policies. Otherwise, \u0026ldquo;knowing you\u0026rdquo; becomes \u0026ldquo;overreading you.\u0026rdquo;\n8. The Ultimate Direction of the Personal Context Path The personal context path will ultimately become the understanding layer of a mature agent system.\nThe understanding layer answers:\nWho is the user? What is the current task? What background is relevant? Which memories are trustworthy? What data can be used? What information should be hidden? What context should be passed to the planning and action layers? Without the understanding layer, no matter how powerful the agent, users still need to repeatedly explain background.\nBut just having the understanding layer isn\u0026rsquo;t enough either. If the agent only understands the user but can\u0026rsquo;t execute tasks or consolidate capabilities, it stays at the knowledge management or personal search tool stage.\nSo the personal context path must ultimately merge with the execution path and the self-evolution path.\n9. Summary The personal context path solves whether an agent can truly understand the user.\nIt will develop from simple preference memory, to workspace context, to multi-source data integration, and finally to a controllable personal memory system.\nThis path has high long-term value because context is harder to migrate than models and closer to real workflows.\nBut its prerequisite is trust.\nUsers won\u0026rsquo;t entrust their email, documents, calendar, chat, code, and task systems to an unexplainable, uneditable, undeletable black box.\nThe next article will discuss the final form after the three paths converge: the governed Agent Runtime.\n","permalink":"/en/posts/agent-evolution-personal-context/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: The core of the personal context path is transforming agents from general assistants into assistants that truly understand your situation. It will evolve from preference memory to workspace context, multi-source data integration, and controllable personal memory systems. The long-term moat isn\u0026rsquo;t connector count, but trustworthy, explainable, editable, deletable context capability.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis is the third article in the Agent Evolution Series.\u003c/p\u003e\n\u003cp\u003eThe first discussed the execution path: how agents move from answering to doing.\nThe second discussed the self-evolution path: how agents move from one-shot assistants to long-term collaborators.\u003c/p\u003e","title":"Agent Evolution Series (3): The Personal Context Path — How Agents Truly Understand You"},{"content":" TL;DR: The core of the self-evolution path is making agents stop treating every interaction like a first meeting. It will progress from in-conversation learning to long-term memory, skills, and evaluation/rollback mechanisms. What truly matters is not remembering more automatically, but consolidating effective experience into verifiable, deletable, rollback-capable capability assets.\nThis is the second article in the Agent Evolution Series.\nThe first article discussed the execution path: how agents move from answering to doing.\nThis article discusses the second path: The Self-Evolution Path.\nThis path asks:\nCan an agent learn from historical tasks, user feedback, and failure experiences, becoming a long-term collaborator that gets better with use?\nIf the execution path solves \u0026ldquo;can it do,\u0026rdquo; the self-evolution path solves \u0026ldquo;can it do it better over time.\u0026rdquo;\nProjects like Hermes Agent represent this path. They emphasize skills, long-term memory, user models, multi-model support, sandbox backends, task trajectory generation and compression, aiming for agents to not just complete tasks but to consolidate experience.\n1. What Is the Self-Evolution Path Many AI assistants today have a clear problem: they can be smart within a conversation, but the next time it\u0026rsquo;s like the first meeting again.\nUsers often need to repeatedly explain:\nWhat this project is; What tech stack is used; What the coding conventions are; Where common tools are located; Which processes should be avoided; Which commands have failed before; What output style the user prefers; What internal team conventions exist. This makes it hard for AI to become a true long-term collaborator.\nThe self-evolution path addresses this problem.\nIt wants agents to accumulate experience from each task, forming long-term capabilities.\nBut \u0026ldquo;self-evolution\u0026rdquo; here shouldn\u0026rsquo;t be understood as models mysteriously awakening, nor should it mean agents can rewrite themselves without limit.\nA more precise definition is:\nThe self-evolution path consolidates effective experience from historical tasks into reusable, evaluable, versionable, deletable, rollback-capable capability assets.\nThis is the engineering-feasible version of self-evolution.\n2. Why the Self-Evolution Path Emerged The fundamental reason is that users don\u0026rsquo;t want to teach the agent from scratch every time.\nA truly valuable agent shouldn\u0026rsquo;t only be smart within the current conversation.\nIf every task requires users to re-explain project background, tool paths, historical conventions, and failure experiences, the agent can never become a long-term collaborator—only a one-shot tool.\nHuman collaborators get better because rapport forms through cooperation. The first time requires a lot of background explanation; after the second and third times, communication costs decrease.\nAgents should be the same.\nThey should gradually learn:\nWhat tasks users commonly do; Which tools are most frequently used; Which workflows have been proven effective; Which errors have occurred before; What should be checked first for certain tasks; Which risky actions must pause for confirmation; Which repetitive processes can be encapsulated as skills. This path develops along \u0026ldquo;in-conversation learning → long-term memory → skills → evaluation and rollback\u0026rdquo; for internal reasons.\nBecause for an agent to get better with use, it must first solve continuity, then reusability, and finally correctness.\nIn-conversation learning only solves the current task; long-term memory solves cross-task continuity; skills turn experience into reusable processes; evaluation and rollback prevent erroneous experience from becoming permanently entrenched.\nSo the core of the self-evolution path is not \u0026ldquo;the more remembered, the better,\u0026rdquo; but:\nConsolidate effective experience while preventing erroneous experience from contaminating future tasks.\nIf an agent can\u0026rsquo;t accumulate experience, it\u0026rsquo;s just a one-shot assistant. If it can accumulate, verify, and correct experience, it can become a long-term collaborator.\n3. Stages of the Self-Evolution Path The self-evolution path also doesn\u0026rsquo;t happen overnight. It roughly goes through four stages.\n1. In-Conversation Learning The earliest learning occurs within a single conversation.\nThe user provides context in the current conversation, and the agent uses this information to complete the task.\nThis already improves the experience somewhat, but the problem is that once the context ends, the information disappears.\nThis capability is more like short-term working memory than long-term learning.\n2. Long-Term Memory The second stage is long-term memory.\nAgents begin saving stable, long-term valuable information: user preferences, project background, team conventions, tool configurations, historical decisions, common problems, and proven effective workflows.\nLong-term memory significantly reduces repetitive explanations from users.\nBut it also brings new problems:\nWhat\u0026rsquo;s worth remembering? When should it be used? What happens when old memories expire? What if old and new memories conflict? How to correct wrong memories? How can users view and delete memories? If these problems aren\u0026rsquo;t solved well, long-term memory turns from an asset into a source of contamination.\n3. Skills Stage The third stage is skills.\nSkills are one of the most critical carriers in the self-evolution path.\nA skill can be understood as an agent\u0026rsquo;s reusable capability unit, potentially containing: task steps, tool descriptions, input/output formats, applicable conditions, examples, checklists, risk warnings, failure handling methods, and evaluation criteria.\nLong-term memory is oriented toward \u0026ldquo;facts and preferences,\u0026rdquo; while skills are oriented toward \u0026ldquo;how to do things.\u0026rdquo;\nFor example:\nHow to review a PR; How to generate a research article; How to release a version; How to handle a certain type of data report; How to run tests in a specific project; How to troubleshoot a certain type of production issue. The value of skills is that they can be file-based, structured, readable, editable, and version-managed.\nThis transforms agent capability consolidation from a black box into a governable asset.\n4. Evaluation and Rollback Stage The fourth stage is evaluation and rollback.\nThis is the key to the self-evolution path truly reaching maturity.\nIf an agent automatically generates many memories and skills without an evaluation mechanism, it may not get better with use—it may get messier.\nA mature system must be able to answer:\nDoes this skill actually improve success rates? Is this memory still valid? Is this experience just a one-time occurrence? Is the new skill better than the old version? If the new version fails, can it be rolled back? Can users delete erroneous memories? Can teams audit skill changes? Only when reaching this stage does self-evolution become more than just a slogan.\n4. The Value and Risk of Long-Term Memory Long-term memory is the foundational capability of the self-evolution path, but it\u0026rsquo;s also dangerous.\nIts value lies in giving agents continuity. Its risk lies in errors continuously affecting future tasks.\n1. Value of Memory Long-term memory allows agents to remember stable information: user habits, project structure, frequently used commands, team rules, historical decisions, verified workflows.\nThis reduces repetitive communication and improves collaboration efficiency.\n2. Risk of Memory But memory can also cause contamination:\nTreating temporary preferences as long-term preferences; Continuing to use outdated rules; Taking erroneous summaries as facts; Generalizing special project experience to all projects; Retrieving sensitive context in irrelevant tasks. Wrong answers usually only affect one task. Wrong memories affect many tasks.\nSo long-term memory must be explainable, editable, deletable, and have conflict resolution mechanisms.\n5. Why Task Trajectory Compression Matters A complex task often generates a long trajectory:\nUser goals; Agent\u0026rsquo;s plan; Tool call records; Materials read; Errors encountered; Attempted solutions; The ultimately successful path; User feedback. These trajectories contain a lot of experience. But it\u0026rsquo;s impossible to stuff all historical trajectories into future context.\nSo compression is needed.\nTask trajectory compression aims to distill future-usable experience from a single task.\nFor example:\nWhich error was the real cause; Which path was proven ineffective; Which flow can be reused; Which constraint must be remembered; Whether a new skill should be generated. But compression also has risks. If compression goes wrong, the agent may treat failure experience as success, coincidental conditions as general rules, or drop critical constraints.\nSo trajectory compression can\u0026rsquo;t just be summarization; it must combine evaluation, user feedback, and version management.\n6. The Biggest Problem of the Self-Evolution Path: Self-Contamination The self-evolution path fears not learning too little, but learning wrong.\nSelf-contamination may manifest as:\nWrong memories used long-term; Failed workflows packaged as skills; Outdated APIs still being called; Temporary preferences mistaken for long-term preferences; Task trajectories incorrectly compressed; Auto-generated skills used without testing. These problems are insidious because they\u0026rsquo;re not one-time errors, but long-term bias.\nSo the real challenge of the self-evolution path is not \u0026ldquo;automatic learning,\u0026rdquo; but \u0026ldquo;governed learning.\u0026rdquo;\n7. What a Mature Self-Evolution System Should Have A mature self-evolution system should have a complete skill and memory lifecycle.\nAt minimum:\nCreation: Generate memories and skills from user instructions, task trajectories, or team workflows; Evaluation: Verify whether they\u0026rsquo;re actually useful; Version Management: Preserve historical changes; Rollback: Recover old versions after errors; Deletion: Remove invalid, outdated, or erroneous content; Audit: Know who created it, when it was updated, and which tasks it affected; Permission Control: Different scenarios can only use necessary memories and skills. This means that in the future, the core moat of self-evolving agents isn\u0026rsquo;t just model capability, but capability asset management capability.\n8. The Ultimate Direction of the Self-Evolution Path The self-evolution path will ultimately become the learning layer of a mature agent system.\nIt connects context and execution. Personal context tells the agent: who the user is, what the project is, and what background the current task has. The execution system calls tools and completes real operations. The self-evolution layer consolidates task experience, making future execution faster, more accurate, and less reliant on user repetition.\nIts final form is not a fully autonomous black-box intelligence, but:\nAn evaluable, versionable, rollback-capable, deletable learning layer.\n9. Summary The self-evolution path solves whether an agent can get better with use.\nIt will develop from in-conversation learning, to long-term memory, to skills, and finally to evaluation and rollback.\nThe key to this path is not having the agent automatically write down more things, but establishing a reliable experience consolidation mechanism.\nA truly mature self-evolving agent must prove that what it learns is correct, useful, and controllable, and can correct and roll back when learning goes wrong.\nThe next article will discuss the third path: the Personal Context Path. It focuses not on whether the agent can do or learn, but on whether the agent can truly understand the user.\n","permalink":"/en/posts/agent-evolution-self-improvement/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: The core of the self-evolution path is making agents stop treating every interaction like a first meeting. It will progress from in-conversation learning to long-term memory, skills, and evaluation/rollback mechanisms. What truly matters is not remembering more automatically, but consolidating effective experience into verifiable, deletable, rollback-capable capability assets.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis is the second article in the Agent Evolution Series.\u003c/p\u003e\n\u003cp\u003eThe first article discussed the execution path: how agents move from answering to doing.\u003c/p\u003e","title":"Agent Evolution Series (2): The Self-Evolution Path — How Agents Get Better with Use"},{"content":" TL;DR: The core of the execution path is moving agents from \u0026ldquo;giving advice\u0026rdquo; to \u0026ldquo;completing tasks.\u0026rdquo; It will evolve from tool calling to browser operations, local automation, and governed action layers. The real moat isn\u0026rsquo;t how many tools you can call, but how many things you can stably accomplish within the constraints of permissions, auditing, and rollback.\nThis is the first article in the Agent Evolution Series.\nThe series is divided into five parts: the first three deconstruct the three main agent paths, the fourth discusses the final form after convergence, and the fifth addresses a more fundamental question: why agents evolve along these paths.\nThis article only discusses the first path: The Execution Path.\nThe execution path asks a very direct question:\nCan an agent not just answer questions, but actually complete tasks?\nIf the previous generation of AI assistants focused on \u0026ldquo;generating answers,\u0026rdquo; the execution path pushes agents into the next phase: from answer systems to action systems.\n1. What Is the Execution Path The core goal of the execution path is moving agents from \u0026ldquo;giving suggestions\u0026rdquo; to \u0026ldquo;doing things.\u0026rdquo;\nTraditional AI assistants typically work like this: users ask questions, AI gives text answers. For example:\nTell me how to fix this error; Write an email for me; Summarize this article; Give me a web automation script; Tell me how to organize these files. These capabilities are valuable, but they still operate at the suggestion level. The actual execution still requires human action.\nThe execution path addresses the next step:\nNot just telling users how to fix code, but directly modifying code and running tests; Not just writing email drafts, but putting them in the drafts folder upon authorization; Not just telling users how to fill a form, but opening the web page, filling the form, and submitting; Not just explaining a command, but running it, checking output, and handling failures; Not just summarizing a task, but creating an issue, assigning an owner, and updating status. So the essence of the execution path is:\nTransforming the agent from an information generator into a task executor.\nProjects like OpenClaw represent this path. They emphasize local execution, multi-tool calling, chat interfaces, browser operations, file access, and workflow automation, attempting to bring agents into the user\u0026rsquo;s actual work environment.\n2. Why the Execution Path Emerged The fundamental reason the execution path emerged is that what users truly need is not \u0026ldquo;answers\u0026rdquo; but \u0026ldquo;results.\u0026rdquo;\nWhen users ask AI to analyze an error, they don\u0026rsquo;t want an explanation—they want the problem fixed. When users ask AI to write an email, they don\u0026rsquo;t want text—they want communication completed. When users ask AI to summarize materials, they don\u0026rsquo;t want a summary itself—they want to advance a decision, writing, or project.\nSo when AI can answer \u0026ldquo;how to do it,\u0026rdquo; users naturally ask the next question:\nSince you know how to do it, why can\u0026rsquo;t you just do it for me?\nThis is why the execution path emerged.\nIt\u0026rsquo;s not about showing off tool calling or making agents seem more robotic; it\u0026rsquo;s about shortening the distance from \u0026ldquo;knowing\u0026rdquo; to \u0026ldquo;completing.\u0026rdquo;\nThe execution path is the first to be felt by users because its value is most direct: as soon as an agent completes a real action, the user immediately sees time saved.\nFor example:\nAutomatically organizing files; Batch web scraping; Modifying code based on requirements; Syncing meeting notes to task systems; Extracting information from email and updating spreadsheets; Running tests, locating failures, attempting fixes. This path develops along \u0026ldquo;tool calling → browser and computer use → local automation → Agent Runtime\u0026rdquo; for good reason.\nBecause for an agent to complete real tasks, it must go through four progressive challenges:\nFirst connect to external tools, otherwise it can only answer; Then operate web pages and GUIs, otherwise it can\u0026rsquo;t enter real systems; Then enter the local environment, otherwise it can\u0026rsquo;t handle files, code, and personal workflows; Finally establish permissions, auditing, and rollback, otherwise it can\u0026rsquo;t earn long-term trust. So the essence of the execution path is not \u0026ldquo;making agents more aggressive,\u0026rdquo; but making agents increasingly controllable in increasingly real environments.\n3. Stages of the Execution Path The execution path doesn\u0026rsquo;t happen overnight. It roughly goes through four stages.\n1. Tool Calling Stage The earliest execution comes from tool calling.\nModels no longer just generate text, but can call external tools: search tools, file read/write tools, database query tools, API calling tools, code execution tools, calendar/email/task system tools.\nThis step connects the model to the outside world. Without tool calling, models can only talk; with tool calling, models start to do.\nBut tool calling is just the starting point. What\u0026rsquo;s really hard is: the agent needs to know when to call a tool, which tool to call, what parameters to pass, how to judge the result, and how to adjust strategy upon failure.\nAn agent supporting many tools doesn\u0026rsquo;t mean it has real execution capability. Real execution comes from stably completing multi-step tasks.\n2. Browser and Computer Use Stage The second stage is browser and computer use.\nIn the real world, many tasks don\u0026rsquo;t have clean APIs and can only be done through GUIs. For example: logging into backend systems, downloading reports, filling in web forms, uploading files, modifying SaaS configurations, operating traditional enterprise software.\nAt this stage, agents need to understand screens, web pages, and interface states, and operate through mouse, keyboard, browser, or system interfaces.\nThis allows agents to enter many real business scenarios. But problems also multiply: page layouts may change, button positions may differ, login states may expire, popups may interrupt flows, external pages may contain malicious prompts, agents may misclick, accidentally delete, or mistakenly submit.\nSo browser automation and computer use are not just about \u0026ldquo;clicking buttons\u0026rdquo;—they\u0026rsquo;re a comprehensive test of perception, planning, error correction, and security boundaries.\n3. Local Automation Stage The third stage is local automation.\nAgents begin accessing user local files, command lines, development environments, and system resources.\nThis significantly boosts productivity because many high-value tasks occur in local environments: modifying code, running tests, analyzing logs, organizing materials, batch-processing files, generating reports, calling local scripts.\nBut local execution also means higher risk. Once an agent can run shell, modify files, access credentials, or call local tools, it must be tightly constrained.\nAt this point, the execution path transitions from \u0026ldquo;automation assistant\u0026rdquo; to \u0026ldquo;security agent.\u0026rdquo;\n4. Agent Runtime Stage Finally, the execution path leads to Agent Runtime.\nAt this point, the agent is no longer just a chatbot that can call tools, but a runtime system.\nIt needs to manage: tool registration, permission levels, operation logs, approval workflows, task queues, credential isolation, sandbox execution, failure retry, human takeover, and rollback mechanisms.\nIn other words, mature execution isn\u0026rsquo;t just \u0026ldquo;can do,\u0026rdquo; but \u0026ldquo;can reliably do within boundaries.\u0026rdquo;\n4. Core Bottlenecks of the Execution Path The biggest bottleneck of the execution path is not whether the model can operate tools, but whether it can operate tools safely, reliably, and controllably.\n1. Reliability Bottleneck Real tasks are usually long-running processes.\nThey contain many state changes: reading information, judging goals, selecting tools, executing operations, checking results, handling failures, and re-planning when necessary.\nIf any step goes wrong, the task may fail. So the execution path must solve the stability of long-horizon tasks, not just the success rate of single tool calls.\n2. Permission Bottleneck Once an agent can operate real systems, it must answer permission questions.\nCan it read files? Can it write files? Can it execute commands? Can it send emails? Can it access databases? Can it modify production systems?\nDifferent tasks require different permissions. Mature agents must practice least privilege, not default to having all capabilities.\n3. Security Bottleneck Execution agents will read web pages, emails, documents, code repos, and search results—all of which may contain malicious instructions.\nWhen agents only answer questions, prompt injection may cause wrong answers. But when agents can call tools, prompt injection may cause: data leakage, unauthorized access, erroneous submissions, dangerous command execution, or critical configuration changes.\nThe stronger the execution capability, the more serious the security problem.\n4. Audit and Rollback Bottleneck Enterprise and high-value personal workflows need to know what the agent did.\nMature agents must record: which steps were executed, which tools were called, which data was read, which files were modified, which actions required user confirmation, which step failed, and whether it can be undone.\nWithout audit and rollback, execution agents will struggle to enter critical workflows.\n5. Signs of Maturity for the Execution Path Once the execution path matures, the evaluation criteria shouldn\u0026rsquo;t be \u0026ldquo;how many tools are connected.\u0026rdquo;\nMore important metrics:\nCan tasks be stably completed? Are permissions granular enough? Do high-risk actions require confirmation? Are operations auditable? Can errors be rolled back? Is external content treated as untrusted input? Are credentials isolated? Can the system recover or hand off to humans after failure? In a word:\nThe competitive moat of the execution path is not how much an agent dares to do, but how many things it can stably accomplish within tight boundaries.\n6. The Ultimate Direction of the Execution Path The execution path will ultimately become the action layer of a mature agent system.\nIt transforms user goals and system plans into real operations. But it won\u0026rsquo;t become the endgame of agents alone.\nThe reason is simple:\nWithout personal context, the agent doesn\u0026rsquo;t know the task background; Without self-evolution, the agent can\u0026rsquo;t accumulate historical experience; Without permission governance, the stronger the execution, the more dangerous. So the final form of the execution path is not \u0026ldquo;an AI that can freely operate a computer,\u0026rdquo; but:\nAn authorizable, auditable, rollback-capable, governable action layer.\nIt will merge with the personal context path and the self-evolution path to become a key component of the mature Agent Runtime.\n7. Summary The execution path solves the question of whether an agent can do things.\nIt will start from tool calling, develop into browser and computer use, then enter local automation, and ultimately become the action layer in Agent Runtime.\nThis path has the greatest short-term value because it most directly saves users time.\nBut its ceiling depends not on the number of tools, but on reliability, permissions, security, auditing, and rollback.\nThe next article will discuss the second path: the Self-Evolution Path. It focuses not on whether an agent can do something once, but on whether an agent can accumulate experience from each task and grow stronger with use.\n","permalink":"/en/posts/agent-evolution-execution-path/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e: The core of the execution path is moving agents from \u0026ldquo;giving advice\u0026rdquo; to \u0026ldquo;completing tasks.\u0026rdquo; It will evolve from tool calling to browser operations, local automation, and governed action layers. The real moat isn\u0026rsquo;t how many tools you can call, but how many things you can stably accomplish within the constraints of permissions, auditing, and rollback.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThis is the first article in the Agent Evolution Series.\u003c/p\u003e\n\u003cp\u003eThe series is divided into five parts: the first three deconstruct the three main agent paths, the fourth discusses the final form after convergence, and the fifth addresses a more fundamental question: why agents evolve along these paths.\u003c/p\u003e","title":"Agent Evolution Series (1): The Execution Path — From Answering to Doing"},{"content":"2026 marks an interesting fork in open-source agents: some projects focus on \u0026ldquo;making AI actually do things for you,\u0026rdquo; others emphasize \u0026ldquo;long-term self-growth,\u0026rdquo; and still others prioritize \u0026ldquo;understanding you before taking action.\u0026rdquo;\nOpenClaw, Hermes Agent, and OpenHuman represent three different approaches:\nOpenClaw: A \u0026ldquo;versatile personal execution agent,\u0026rdquo; strong in chat interfaces, automation, system control, and multi-platform integration. Hermes Agent: A \u0026ldquo;self-improving developer/research agent,\u0026rdquo; strong in skill learning, sandbox backends, tool systems, and long-running operation. OpenHuman: A \u0026ldquo;local-first personal memory agent,\u0026rdquo; strong in understanding users, syncing personal data, memory trees, and low-barrier desktop experiences. 1. OpenClaw: The Personal Automation Agent That Actually Does Things OpenClaw positions itself as a Personal AI Assistant, with the core slogan: \u0026ldquo;The AI that actually does things.\u0026rdquo;\nIt\u0026rsquo;s not just a chatbot, but a local-first agent running on the user\u0026rsquo;s own device. Users can send commands through WhatsApp, Telegram, Discord, Slack, Signal, iMessage, and other chat interfaces, asking it to handle email, calendar, files, browsers, scripts, code, reminders, web tasks, and more.\nOpenClaw\u0026rsquo;s Strengths 1. Multi-Chat Interface Strength A key feature of OpenClaw is that it doesn\u0026rsquo;t require users to migrate to a new app—it plugs into the chat tools you already use.\nFor example, you can ask it to execute tasks through Telegram or Slack, just like messaging an assistant.\n2. Strong Local Execution OpenClaw can read files, run shells, control browsers, call APIs, manage sessions, and extend capabilities through skills and plugins.\nThis makes it well-suited for \u0026ldquo;real-world automation,\u0026rdquo; not just answering questions.\n3. Broad Integration Public materials mention support for Gmail, GitHub, Spotify, Obsidian, browsers, Claude, GPT, and numerous chat platforms.\nThe GitHub README lists WhatsApp, Telegram, Slack, Discord, Google Chat, Signal, iMessage, IRC, Teams, Matrix, Feishu, LINE, Mattermost, WeChat, QQ, and more.\n4. High Community and Ecosystem Activity Based on public data, OpenClaw has high GitHub engagement, fork counts, issues, and PRs, indicating a relatively large open-source community.\nOpenClaw\u0026rsquo;s Weaknesses and Risks 1. High Permissions, Prominent Security Risks OpenClaw\u0026rsquo;s strength is also where its risk lies. It can access local files, browsers, shell, email, payment, or third-party services.\nIf affected by prompt injection, malicious web pages, malicious messages, or misconfiguration, the risks are much greater than with ordinary chatbots.\n2. Enterprise Deployment Requires Governance If employees privately connect OpenClaw to company email, GitHub, Slack, or file systems, it creates \u0026ldquo;shadow IT.\u0026rdquo;\nEnterprises need auditing, permission isolation, sandboxes, logging, approval workflows, and data compliance mechanisms.\n3. Non-Technical Users May Find Configuration Difficult Although positioned as a personal assistant, many powerful features require understanding concepts like gateway, daemon, sandbox, session, channel, and tool policy.\nOrdinary users wanting \u0026ldquo;out-of-the-box\u0026rdquo; experience may find it complex.\nOpenClaw\u0026rsquo;s Ideal Use Cases Personal automation assistant Email, calendar, file, browser task processing Developer daily workflow automation Multi-chat-platform unified AI assistant Persistent agent on personal servers or local devices Experimental automation for technical teams Who Should Use OpenClaw? Best for:\nTechnically capable individual users Developers, indie hackers, automation enthusiasts Those who want AI integrated into WhatsApp, Slack, Telegram Those who understand permission risks and are willing to configure sandboxes Less suitable for:\nCompletely non-technical users Those with no concept of privacy and permission configuration Enterprises without security governance capabilities for large-scale deployment 2. Hermes Agent: The Self-Improving Agent That Accumulates Skills Hermes Agent comes from NousResearch/hermes-agent, positioned as \u0026ldquo;The agent that grows with you.\u0026rdquo;\nIt\u0026rsquo;s more of a long-running agent for developers, researchers, and heavy automation users. It emphasizes self-improvement, long-term memory, user modeling, skill generation, multi-model support, tool gateways, sandbox backends, and multi-interface interaction.\nHermes Agent\u0026rsquo;s Strengths 1. Self-Improvement and Skill System Hermes Agent doesn\u0026rsquo;t just execute tasks; it emphasizes generating and improving skills from task experience.\nThis makes it more like a long-term collaborator, not a one-shot tool caller.\n2. Rich Sandbox and Runtime Backends Public materials show support for local, Docker, SSH, Singularity, Modal, Daytona, Vercel Sandbox, and other backends.\nThis is valuable for developers, researchers, and those needing isolated execution environments.\n3. Strong Multi-Model and Multi-Supplier Support Hermes Agent supports Nous Portal, OpenRouter, NovitaAI, NVIDIA NIM, OpenAI, Hugging Face, AWS Bedrock, LM Studio, Azure AI Foundry, custom endpoints, and more.\nThis means it\u0026rsquo;s better suited for users who enjoy tinkering with model routing, costs, performance, and local models.\n4. Developer Experience It supports CLI/TUI, cron, subagents, MCP, tool gateways, trajectory generation, trajectory compression, and more.\nThese capabilities are clearly more oriented toward developers, researchers, and agent engineering users.\nHermes Agent\u0026rsquo;s Weaknesses and Risks 1. High Learning Curve Although it has a one-line install command, its capability system is complex: models, tools, sandboxes, skills, MCP, gateways, cron, subagents.\nOrdinary users may not know where to start.\n2. Higher Infrastructure Requirements Hermes Agent is better suited for long-running environments like servers, VPS, local dev machines, or GPU environments.\nIf you just want a simple desktop assistant, it may feel \u0026ldquo;too heavy.\u0026rdquo;\n3. Windows Native Support Needs Caution Documentation indicates Windows native support has beta coloring; production or stable use recommends WSL2.\n4. Self-Improvement Also Needs Governance Agents automatically generating skills, updating skills, and long-term memorization of user habits also require review mechanisms.\nOtherwise, erroneous skills, outdated preferences, or contaminated memory may affect future tasks.\nHermes Agent\u0026rsquo;s Ideal Use Cases Developer long-running agent Research agent experiments Automated task scheduling Long-running assistant on servers/VPS Multi-model, multi-tool, multi-sandbox environments Agent skill generation, reuse, evaluation Complex workflows requiring subagents for parallel task processing Who Should Use Hermes Agent? Best for:\nAI agent developers Researchers DevOps and platform engineers Advanced automation users Teams wanting to research self-improving agents Those with experience using servers, containers, MCP, model APIs Less suitable for:\nThose who just want a quick personal desktop AI assistant Those who don\u0026rsquo;t want to configure models and toolchains Users without basic command-line experience 3. OpenHuman: The Local Memory Agent That Understands You First OpenHuman comes from tinyhumansai/openhuman, positioned more as a local-first personal AI super intelligence.\nIts core selling point is not \u0026ldquo;how many tools it can connect,\u0026rdquo; but \u0026ldquo;understanding your personal context first.\u0026rdquo;\nIt emphasizes local memory, privacy, desktop experience, 118+ third-party integrations, Memory Tree, Obsidian Wiki, TokenJuice compression, and data source syncing from Gmail, Notion, GitHub, Slack, Stripe, Calendar, Drive, Linear, Jira, and more.\nOpenHuman\u0026rsquo;s Strengths 1. Memory System OpenHuman\u0026rsquo;s Memory Tree compresses and hierarchically summarizes user data, storing it in local SQLite.\nIt also writes knowledge as Obsidian-style Markdown files, making it easy for users to view, edit, and migrate.\nThis is more transparent than \u0026ldquo;black-box embedding memory\u0026rdquo; and better suited for those who value personal knowledge management.\n2. Local-First and Privacy Narrative It emphasizes personal data, local models, local memory, and user control.\nFor many concerned about cloud AI assistants reading private data, this direction is very appealing.\n3. More User-Friendly Experience Compared to OpenClaw and Hermes Agent, OpenHuman feels more like a desktop product.\nIt has an official website with downloadable installers, plus desktop UI, voice, meeting participation, search, scraping, coding tools, and more.\n4. Strong Personal Data Integration 118+ integrations are a key selling point.\nIt attempts to unify Gmail, Notion, GitHub, Slack, Calendar, Drive, Linear, Jira, and other personal or work data into a long-term context understandable by an agent.\nOpenHuman\u0026rsquo;s Weaknesses and Risks 1. Still Early Beta The GitHub README explicitly labels it as Early Beta, noting it\u0026rsquo;s still under rapid development.\nThis means stability, compatibility, installation experience, and security boundaries may not be mature.\n2. More Connections, More Concentrated Risk OpenHuman\u0026rsquo;s strength is \u0026ldquo;knowing a lot about you,\u0026rdquo; but that\u0026rsquo;s also its biggest risk.\nIf it connects email, code repos, calendar, chat history, and payment tools, then aggregates everything into local SQLite and Markdown, a local breach could expose highly concentrated information.\n3. Default Experience May Still Depend on Managed Backend Although it emphasizes local-first, public materials show that login, model routing, search agents, integration/OAuth flows, and other managed experiences may still depend on the OpenHuman backend or Composio.\nIf users want \u0026ldquo;fully local,\u0026rdquo; they need additional configuration for models, search, Composio credentials, etc.\n4. License and Commercial Use OpenHuman\u0026rsquo;s repository is GPL-3.0 licensed.\nIf enterprises want to build on it or integrate it into internal products, they need to carefully evaluate the GPL license implications.\nOpenHuman\u0026rsquo;s Ideal Use Cases Personal knowledge management Private AI assistant Local memory base Context integration across email, calendar, docs, code, task systems AI-powered knowledge management for Obsidian users Personal users wanting AI to understand them long-term Lightweight team or personal productivity experiments Who Should Use OpenHuman? Best for:\nThose who value personal memory and knowledge management Heavy users of Obsidian, Notion, Gmail, GitHub, Slack Those wanting a desktop, local-first AI assistant Users who don\u0026rsquo;t want to start from the command line Early adopters willing to accept beta product instability Less suitable for:\nEnterprise production environments requiring high stability Those unwilling to grant access to sensitive data like email, calendar, code repos Those who cannot accept the risk of centralized local data storage Teams needing permissive commercial licenses for secondary development 4. Side-by-Side Comparison Dimension OpenClaw Hermes Agent OpenHuman Core Positioning Local-first personal execution agent Self-improving developer/research agent Local-first personal memory agent Keywords Chat interface, automation, browser, shell, files, plugins Skills, self-improvement, sandbox, MCP, multi-model, long-running Memory Tree, Obsidian, local memory, 118+ integrations, desktop Learning Curve Medium-high High Medium Target Users Technical individuals, automation enthusiasts, developers Agent developers, researchers, platform engineers Personal productivity users, knowledge management, privacy-conscious Strongest Capability Actually executing real-world tasks Continuous learning and tool-based workflows Quickly understanding user context Automation Very strong Very strong, but more engineering-oriented Medium-strong, more personal data and knowledge flow Memory Long-term memory Long-term memory and user modeling Memory is core product capability Integration Chat platforms, browser, files, shell, plugins CLI/TUI, MCP, messaging platforms, tool gateway, sandbox Gmail, Notion, GitHub, Slack, Calendar, Drive, Linear, Jira Security Risk High (due to broad permissions) Medium-high (due to complex tools and self-improvement) High (due to aggregating large amounts of personal data) Enterprise Use Usable but needs strong governance Better for R\u0026amp;D/platform internal experiments Caution needed, review privacy and GPL License MIT MIT GPL-3.0 Maturity High ecosystem activity, security governance pressure Strong engineering, suitable for professional users Early Beta, good experience direction but stability needs observation 5. How to Choose If You Want \u0026ldquo;AI to Actually Do Things for Me\u0026rdquo; Choose OpenClaw.\nIt\u0026rsquo;s ideal for connecting AI to chat tools to handle email, calendar, files, browsers, scripts, code, messages, and more.\nBut be prepared to accept and manage its high-permission risks.\nIf You Want to Research \u0026ldquo;Growing Agents\u0026rdquo; Choose Hermes Agent.\nIt\u0026rsquo;s ideal for agent engineering, skill learning, self-improvement, multi-model, multi-tool, multi-sandbox, and long-running experiments.\nIt\u0026rsquo;s not the lightest personal assistant, but great for technical teams and researchers.\nIf You Want \u0026ldquo;AI to Understand Me First\u0026rdquo; Choose OpenHuman.\nIt\u0026rsquo;s ideal for integrating personal data, knowledge bases, email, calendar, code, and documents into long-term memory.\nBut it\u0026rsquo;s still Early Beta, and the privacy risks of centralized data must be carefully evaluated.\n6. Conclusion OpenClaw, Hermes Agent, and OpenHuman represent the next generation of open-source agents, but they solve different problems.\nOpenClaw is more like a \u0026ldquo;versatile personal assistant that can actually do things.\u0026rdquo; Hermes Agent is more like a \u0026ldquo;developer/research agent that accumulates skills.\u0026rdquo; OpenHuman is more like a \u0026ldquo;local personal memory system that understands you before helping you.\u0026rdquo;\nLooking at the evolution of mature agents, these three represent three paths:\nExecution Path: OpenClaw Self-Evolution Path: Hermes Agent Personal Context Path: OpenHuman The truly powerful personal agent of the future will likely combine all three capabilities: understanding you, learning long-term, and safely executing real tasks on your behalf.\n","permalink":"/en/posts/openclaw-hermes-openhuman-comparison/","summary":"\u003cp\u003e2026 marks an interesting fork in open-source agents: some projects focus on \u0026ldquo;making AI actually do things for you,\u0026rdquo; others emphasize \u0026ldquo;long-term self-growth,\u0026rdquo; and still others prioritize \u0026ldquo;understanding you before taking action.\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eOpenClaw, Hermes Agent, and OpenHuman represent three different approaches:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eOpenClaw\u003c/strong\u003e: A \u0026ldquo;versatile personal execution agent,\u0026rdquo; strong in chat interfaces, automation, system control, and multi-platform integration.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eHermes Agent\u003c/strong\u003e: A \u0026ldquo;self-improving developer/research agent,\u0026rdquo; strong in skill learning, sandbox backends, tool systems, and long-running operation.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpenHuman\u003c/strong\u003e: A \u0026ldquo;local-first personal memory agent,\u0026rdquo; strong in understanding users, syncing personal data, memory trees, and low-barrier desktop experiences.\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n\u003ch2 id=\"1-openclaw-the-personal-automation-agent-that-actually-does-things\"\u003e1. OpenClaw: The Personal Automation Agent That Actually Does Things\u003c/h2\u003e\n\u003cp\u003eOpenClaw positions itself as a \u003cstrong\u003ePersonal AI Assistant\u003c/strong\u003e, with the core slogan: \u003cstrong\u003e\u0026ldquo;The AI that actually does things.\u0026rdquo;\u003c/strong\u003e\u003c/p\u003e","title":"OpenClaw vs Hermes Agent vs OpenHuman: Which Open-Source Agent Is Right for You?"},{"content":" TL;DR\nThis article is not about \u0026ldquo;which API is cheaper,\u0026rdquo; but about AI Tokens becoming the settlement unit for model capacity—and who will manage the gateway, routing, ledger, quality, and responsibility. Core judgment: Model reselling is not the endgame; model capacity operators are the long-term direction. Gray-market reselling will recede, but unified APIs, multi-model routing, semantic caching, compliance auditing, enterprise governance, and cost optimization will remain. The future competition will not be about \u0026ldquo;how many models you can access,\u0026rdquo; but whether you can operate model capacity into stable, trustworthy, deliverable, auditable, billable, and accountable results.\nIntroduction: Don\u0026rsquo;t Just See Model Reselling as a Cheap API Business Recently, model proxies, API relays, token agents, and unified API platforms have suddenly become very popular.\nAt first glance, many people see it as a \u0026ldquo;cheap API business\u0026rdquo;: someone bundles OpenAI, Claude, Gemini, DeepSeek, Qwen, Kimi and other models behind a single interface, then sells access at a lower price. Users don\u0026rsquo;t need to register on multiple platforms, manage dozens of API keys, or study each model\u0026rsquo;s pricing rules—just top up one balance and call many models.\nOn the surface, this looks like gray-market reselling, shell forwarding, and low-price arbitrage.\nBut if you only see it as a short-term cheap business, you underestimate the industrial shift behind it.\nMy judgment is: The chaos of model proxies, API relays, and token agents is the early-stage disorder of the market after AI Tokens became the settlement unit for model capacity. Gray-market reselling will recede, but unified gateways, multi-model scheduling, unified billing, cost governance, compliance auditing, and capability packaging will not disappear. They will continue to evolve and eventually give birth to a new industrial role: the Model Capacity Operator.\nA Model Capacity Operator is not a simple API reseller, nor does it necessarily own the strongest models. What it truly operates is the gateway, metering, routing, ledger, quality, compliance, and responsibility boundaries of model capacity.\nNext, the article follows a central thread: first, where today\u0026rsquo;s chaos comes from; then, why AI Tokens become the underlying settlement unit; followed by seven analogies to understand its future trends; and finally, why Model Capacity Operators will emerge, how they will evolve, and how different roles should respond.\nYou can think of the rest as nine consecutive questions:\nWhat is the current chaos of model proxies, API relays, and token agents? Why did they emerge, and what business models and gray/black risks lie behind them? Why is AI Token becoming the underlying settlement unit for model capacity? How can seven resource analogies help understand the future of AI Tokens? What do these seven analogies together point to regarding Model Capacity Operators? What are the capability boundaries, business models, and reasons for the emergence of Model Capacity Operators? What development trends will model capacity operations face next? How should individual users, developers, enterprises, and entrepreneurs respond? Finally, how should we judge this industry trajectory? 1. Current State: The Chaos of Model Proxies, API Relays, and Token Agents The most obvious feature of today\u0026rsquo;s model relay market is: demand is real, but forms are chaotic.\nOn one side, users really need a unified gateway.\nIndividual users want to spend less, developers want fewer SDKs to integrate, and enterprises want unified management of model calls. Everyone wants one account, one balance, one API, and one bill to call multiple models simultaneously.\nOn the other side, the supply side is very rough.\nMany model proxy platforms claim to support dozens or even hundreds of models, with prices much lower than official ones, interfaces compatible with OpenAI format, and immediate availability after top-up. They solve some real pain points, but also bring many problems:\nUpstream model sources are opaque; Whether authorization is obtained is unclear; Whether gray key pools are used is unclear; Whether shared accounts, promo arbitrage, or stolen credits are involved is unclear; Whether the model users call is actually the real model is unclear; Whether expensive models are replaced by cheaper ones is unclear; How many layers of proxy the request passes through is unclear; Whether data is recorded, stored, trained on, or reused is unclear; Whether the platform has SLA, audit, compliance, and responsibility commitments is unclear. This is the most typical chaos of today\u0026rsquo;s model proxies, API relays, and token agents:\nWhat users see: cheap, convenient, many models.\nWhat may be behind the platform: gray key pools, unauthorized resale, fake models, black-box pipelines, and security risks.\nThe \u0026ldquo;fake model\u0026rdquo; problem, in particular, will become increasingly serious.\nUsers see the interface returning gpt-4, claude, gemini, or some expensive model name, but whether it\u0026rsquo;s actually that model is hard for ordinary users to verify. Some platforms may use cheap models to impersonate expensive ones, use distilled models to impersonate original ones, use local models to impersonate official ones, or even dynamically downgrade during peak hours.\nThis kind of risk is not simply about \u0026ldquo;a bit more expensive or cheaper\u0026rdquo;—it\u0026rsquo;s about the authenticity of model capability.\nIf users are just using it for casual chat, the problem is not serious. But once it enters code generation, contract review, financial analysis, enterprise knowledge bases, automated agents, or internal tooling, this black-box relay layer becomes a supply chain risk point.\nSo today\u0026rsquo;s chaos can be summed up in one sentence:\nDemand has already emerged, but the market is still in an early, rough, low-trust relay form.\nThis is not the endgame, but a transitional phase.\n2. Why Did It Emerge: Demand, Business Models, and Gray/Black Risks Converging Why did model relays, API relays, and token agents emerge? It can\u0026rsquo;t be explained by \u0026ldquo;people being cheap\u0026rdquo; alone.\nBehind it are at least three forces: real demand, commercial arbitrage, and gray/black risks.\n1. Real Demand: Too Many Models, Too Fragmented Interfaces Today, a developer or enterprise wanting to use large models faces not just one supplier, but a long list: OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Qwen, Kimi, Zhipu, MiniMax, Volcano Ark, Alibaba Cloud Bailian, Tencent Cloud, Huawei Cloud, AWS Bedrock, Azure AI Foundry, Google Vertex AI, plus various open-source and self-hosted models.\nEach platform has its own API format, authentication method, model names, pricing, context length, rate limits, error codes, SDKs, data policies, and compliance requirements.\nIf you\u0026rsquo;re just experimenting individually, this complexity is tolerable. But once you enter development and production environments, the problems become very real:\nShould a product integrate multiple models at once? If the primary model fails, should it automatically switch to a backup? If a task is simple, can it automatically use a cheaper model? If a user uploads sensitive data, should it use a private or trusted cloud model? If a department exceeds budget, can it be rate-limited? If the boss asks how much the company spent on AI this month, can you give a clear answer? So what users truly need is not the \u0026ldquo;relay\u0026rdquo; form itself, but:\nOne account, one balance, one API, one bill, one routing strategy.\nIn the early single-model era, the problem was simple: get a key, write a prompt, send a request, get a result.\nThe multi-model era is different.\nDifferent models suit different tasks. Simple Q\u0026amp;A can use cheap, fast models; translation, summarization, rewriting can use cost-effective general models; complex reasoning needs stronger models; code tasks need models strong at coding; long-document processing needs long-context models; images, audio, video need multimodal models; enterprise private data may need self-hosted models; legal, medical, financial scenarios need vertical industry models.\nSo the question is no longer:\nWhich model is the best?\nBut rather:\nWhich model is most suitable for this current task?\nThat\u0026rsquo;s the value of model routing and unified APIs.\nGoing deeper, AI workloads are different from traditional web APIs. Traditional API requests usually have predictable costs, short responses, and light state; but model requests can consume large amounts of context at once, return long streaming responses, and encounter rate limits, queuing, 429 errors, latency jitter, and service degradation under high concurrency. More troublesome is that underlying model capabilities are highly heterogeneous: the same question given to different models yields different costs, speeds, quality, compliance risks, and responsibility boundaries.\nThis forces the relay layer to upgrade from a \u0026ldquo;passive forwarding channel\u0026rdquo; to an \u0026ldquo;active management pipeline\u0026rdquo;: inbound needs identity, permissions, budget, sensitive data, and prompt injection checks; outbound needs content safety, logging, auditing, watermarking, quality observation, and responsibility attribution. In other words, production-grade AI infrastructure naturally pushes simple API relays toward more complex model capacity operations.\n2. Current Business Models: Low-Price Arbitrage, Unified Gateway, and Balance Pools Today, many model relay platforms have relatively early-stage business models, mainly in several categories:\nThe first is low-price arbitrage. Platforms obtain lower upstream costs through bulk purchasing, promotional credits, regional price differences, account systems, or other means, then sell to users at prices lower than official but higher than their own costs.\nThe second is prepaid balances. Users top up first, and the platform deducts based on token consumption. Simple for users, forming a balance pool and cash flow for the platform.\nThe third is unified API. Platforms package multiple models into a unified interface, allowing developers to call multiple models with one format.\nThe fourth is multi-model packages. Platforms bundle different models into memberships, credit packs, monthly packs, and team packs, so users don\u0026rsquo;t have to study price tables one by one.\nThe fifth is enterprise usage packs. Targeting teams or enterprises, providing budgets, billing, sub-accounts, statistics, rate limiting, logging, and other basic capabilities.\nThese models themselves are not necessarily problematic. The problem is: if the platform has no authorization, no transparent upstream, no security commitments, no audit capabilities, no content watermarking, and no responsibility boundaries, it easily slides into gray or even black chains.\n3. Non-Compliance, Risks, and Gray/Black Factors The risks of the model relay layer are mainly concentrated in several areas.\nFirst, unauthorized resale. Many model service terms do not allow users to resell API keys. If a platform doesn\u0026rsquo;t have upstream authorization, it may essentially be doing unauthorized distribution.\nSecond, gray key pools. Platforms may pool a large number of personal accounts, student credits, trial credits, promotional credits, stolen accounts, or keys of unknown origin, then sell them externally. Such models may have very low prices, but stability, legality, and security are all poor.\nThird, fake models and model downgrading. Platforms may claim to call an expensive model but actually use a cheap model, distilled model, or other substitute. Users have a hard time proving whether they\u0026rsquo;re getting the target model.\nFourth, data leakage. Users may send code, contracts, customer information, enterprise data, database queries, internal API parameters, and agent tool call traces to the relay station. If the relay station has no clear data processing agreement, no security commitment, and no audit mechanism, the risk is very high.\nFifth, Shadow AI. Employees privately purchase API keys, upload corporate data, connect external agents, and link business systems to untrusted models, creating invisible, unmanageable, unauditable AI usage chains within the organization.\nSixth, the Agent era will amplify supply chain risks. In the past, relay stations mainly saw prompts. In the future, agents will call tools, read files, access databases, connect to MCP servers, call internal APIs, and even process credentials and business system results. At this point, the relay layer is no longer just a \u0026ldquo;text forwarder\u0026rdquo; but could become a supply chain attack point.\nSeventh, regulation will increasingly focus on model sources, content watermarking, responsibility for generated content, data cross-border transfers, security assessments, algorithm filing, and call chain auditing.\nWhen model sources need to be explained, generated content needs to be watermarked, and call chains need to be audited, black-box relays will find it hard to survive long-term.\nSo my judgment is:\nGray-market relays will be compressed, but model aggregation will not disappear.\nWhat disappears is opaque, unaccountable, unauditable black-box relays.\nWhat remains is unified gateways, multi-model scheduling, unified billing, enterprise governance, cost optimization, compliance auditing, and industry capability packaging.\n3. Underlying Factor: AI Token Is Becoming the Settlement Unit of the AI Era Why do model relays revolve around tokens? Why must all API platforms eventually deal with billing, settlement, packages, cost attribution, and usage management?\nThe underlying reason is: AI Token is becoming the most important settlement unit of the AI era.\nIt was originally a technical concept in NLP, representing the basic fragment of text processed by a model. But in the commercial era of large models, AI Token is no longer just a text slice—it\u0026rsquo;s the common unit connecting model capability, user tasks, supplier costs, and platform ledgers.\nIn the past, we might only care about input tokens and output tokens. Now we also need to consider cached input, reasoning tokens, long context tokens, vision tokens, audio tokens, video tokens, tool-use tokens, and agent step tokens.\nFor the same 1 million AI Tokens, different models can have vastly different prices; within the same model, input and output prices differ; cache hit and miss prices differ; normal and reasoning modes differ; long context, multimodal, low latency, high reliability all change the price.\nMore precisely, AI Token is evolving from a technical measurement unit into a \u0026ldquo;cognitive load\u0026rdquo; measurement unit in AI-native services.\nIn the traditional internet, users care about bandwidth, latency, and packet loss. In AI services, new experience metrics become:\nTime-to-First-Token (TTFT) Tokens per second (TPS) Context length Cache hit rate Cost per Token Tokens per Watt Task success rate Cross-agent settlement accuracy Goodput: effective output meeting latency SLO This means that in the future, measuring whether an AI service is good depends not just on \u0026ldquo;whether the interface works,\u0026rdquo; but on whether it can stably, with low latency and low cost, generate sufficiently high-quality AI Tokens and ultimately complete user tasks.\nBut here\u0026rsquo;s an important boundary: users don\u0026rsquo;t actually care about AI Tokens themselves.\nWhat users care about is:\nCan this code be written well? Can this document be summarized? Can this contract be reviewed? Can this customer service issue be resolved? Can this agent run through tasks stably? Can this professional advice be verified and held accountable? AI Token is the underlying consumption unit, but what users ultimately buy is the result.\nJust as ordinary users don\u0026rsquo;t care how many data packets are used behind a video call—they only care if the call is smooth—future AI users won\u0026rsquo;t constantly worry about how many tokens they\u0026rsquo;ve consumed, but rather: how much does a task cost? Is the result good? Is it stable? Who\u0026rsquo;s responsible if something goes wrong?\nSo the industrial significance of AI Token is:\nFor model vendors, it\u0026rsquo;s the unit of inference cost.\nFor platforms, it\u0026rsquo;s the unit of ledger and settlement.\nFor enterprises, it\u0026rsquo;s the unit of cost governance.\nFor users, it will ultimately be packaged into plans, task packs, and outcome-based pricing.\nThis is the foundation for all the analogies that follow.\nHowever, AI Token cannot simply be equated to traffic, compute, or electricity.\nThe most important sentence is:\nMobile data transfers information; AI Tokens generate judgments.\nBehind AI Tokens are not ordinary data packets, but answers, code, plans, suggestions, judgments, and even action instructions. Once they enter enterprise workflows, financial risk control, legal review, medical assistance, and automated agents, it\u0026rsquo;s not just about \u0026ldquo;whether the call succeeded,\u0026rdquo; but \u0026ldquo;whether the result is reliable, explainable, auditable, and accountable.\u0026rdquo;\n4. Future Trends of AI Tokens: A Framework of Seven Analogies This section has the highest information density in the entire article. To avoid reader fatigue from diving straight into lengthy analysis, here\u0026rsquo;s a diagram laying out the seven analogies first:\nHere\u0026rsquo;s a summary table for quick reference:\nSummary Table of Seven Analogies for AI Token Future Trends # Analogy Corresponding Trend What to Remember 1 Mobile Data Plans, credit packs, task packs Users won\u0026rsquo;t just care about unit price, but about \u0026ldquo;is it enough, will I exceed, can I share\u0026rdquo; 2 Cloud Computing Cost governance, budget attribution, elastic scaling Enterprises will manage AI Token costs like they manage cloud costs 3 Electricity Stable supply, redundancy, SLA/SLO Model capability will go from \u0026ldquo;can call\u0026rdquo; to \u0026ldquo;must be stably supplied\u0026rdquo; 4 Payment Clearing Network Cross-model, cross-supplier, cross-agent settlement Multi-model era needs ledgers, reconciliation, profit sharing, and dispute resolution 5 CDN Routing, caching, fallback Requests will be dynamically distributed to the most suitable model, not always hitting a single model 6 Enterprise Governance Gateway Permissions, audit, risk control, compliance Enterprises need to know who is using which model, spending how much, and where data goes 7 Professional Services Tiering, verification, responsibility boundaries General tokens will be cheap; professional tokens will command a premium due to result quality and accountability How to read this table: These 7 analogies don\u0026rsquo;t replace each other; they explain 7 different facets of AI Tokens. Together, they point to the core role in the following section: the Model Capacity Operator.\nTo understand the future of AI Tokens, I think seven analogies are useful.\nThese analogies are not meant to say AI Token is completely equivalent to some old resource, but to explain its different trends on the user side, cost side, infrastructure side, settlement side, scheduling side, organizational governance side, and professional service side.\n1. Like Mobile Data: Plans From the user\u0026rsquo;s perspective, AI Token is most like mobile data.\nThe telecom industry once shifted from \u0026ldquo;charging by minutes and SMS\u0026rdquo; to \u0026ldquo;charging by data plans.\u0026rdquo; Behind this was the shift from circuit-switched to packet-switched networks, with the unit of measurement changing from connection duration to data packets and GB.\nThe AI industry is undergoing a similar transition: applications are no longer just charging by software seat or subscription period, but increasingly organizing business models around AI Token consumption, Time-to-First-Token, Token throughput, and task results.\nMobile data has gone through: Per MB billing → Monthly plans → Large data plans → Unlimited plans → Family sharing / Enterprise dedicated lines\nAI Tokens may also go through: Per million tokens → Monthly token packs → Team shared credits → Agent call packs → AI office suites → Industry task packs\nOrdinary users won\u0026rsquo;t care about the price per million AI Tokens for long. They care more about: Is my plan enough? Can this task be completed? Is the result stable? What happens if I go over?\nSo on the user side, the first thing future platforms must do is translate complex model calls into understandable credits, plans, balances, task packs, and outcome-based prices.\nBut the boundary is also clear:\nMobile data transfers information; AI Tokens generate judgments.\nMobile data mainly solves connectivity; AI Tokens also affect answers, decisions, transactions, code, contracts, organizational processes, and automated actions.\n2. Like Cloud Computing: Resourceization and Cost Governance In front of users, AI Token is like data; in the ledgers of platforms and enterprises, AI Token is more like cloud resources.\nCloud computing features on-demand, elastic, resource pooling, pay-per-use, observable, and optimizable.\nAI Token will enter a similar cost governance system:\nWhich department used how many AI Tokens? Which project burns the most money? Which tasks are suitable for caching? Which tasks can use cheap models? Which calls need strong models? What\u0026rsquo;s the cost per API call? What\u0026rsquo;s the AI cost per customer, per order, per task? In the future, enterprises will do AI cost governance the same way they do cloud cost governance.\nThis is why Cost per Token becomes increasingly important. In the past, when enterprises purchased compute, they looked at GPU models, FLOPS, VRAM, and rental unit price; but what large model inference actually delivers to business is not FLOPS, but usable AI Tokens.\nA more expensive new piece of hardware, if it generates more AI Tokens per second, more AI Tokens per watt, and lower cost per million AI Tokens, might actually be the cheaper choice.\nFrom this perspective, data centers will increasingly become \u0026ldquo;AI Token factories\u0026rdquo;: raw materials are electricity, chips, model weights, and data; output is intelligent tokens consumable by applications.\nA platform\u0026rsquo;s cost advantage isn\u0026rsquo;t just buying cheap APIs, but combining underlying compute, caching, batch processing, local models, private models, and cloud models into the lowest unit AI Token cost.\nWe\u0026rsquo;ll also see \u0026ldquo;Model-as-a-Service\u0026rdquo; architectures where high-frequency, routine, sensitive, or low-complexity inference tasks can be offloaded to local or private cloud models, while complex, low-frequency, heavy reasoning tasks go to cloud large models. Otherwise, if enterprises put all automated workflows on public cloud token-metered interfaces, the efficiency gains from AI may be eaten up by continuously growing inference bills, creating structural profit leakage.\n3. Like Electricity: Stable Supply and Infrastructuralization When AI penetrates office work, customer service, R\u0026amp;D, finance, healthcare, and government, model capability will become infrastructure, like electricity.\nToday\u0026rsquo;s software systems can function without AI, but in the future, many software systems, employees, devices, and workflows may continuously call models.\nAt that point, users will care about:\nIs it stable? Will it go down? Is there an SLA? Is there a backup model? Will it slow down during peak hours? Can responsibility be assigned if something goes wrong? The platform\u0026rsquo;s value will upgrade from \u0026ldquo;low-price forwarding\u0026rdquo; to \u0026ldquo;stable supply of intelligent resources.\u0026rdquo;\nGoing further, model capability platforms may participate in building a distributed intelligent infrastructure similar to an AI Grid. The heaviest training and complex reasoning remain in centralized AI factories; regional computing centers handle city-level and industry-level inference loads; edge models on enterprise premises, base stations, and terminal devices handle local high-frequency tasks and personalized context.\nThe purpose is not to sound good conceptually, but to reduce latency, reduce redundant transmission, meet data residency requirements, and extend AI Token generation from a single cloud center to a more distributed intelligent network.\nA point worth keeping from the reference material is that centralized cloud architecture will encounter physical and environmental limits. AI models are energy-intensive workloads, and power supply, chip shortages, cooling bottlenecks, and geopolitical factors will all constrain the infinite expansion of a single centralized cloud. A more sensible future form is not to route all data back to a central large model, but to push some \u0026ldquo;intelligent computing power\u0026rdquo; closer to where data resides: centralized AI factories handle the heaviest training and complex reasoning; regional computing centers handle daily high-intensity inference; small language models on enterprise premises, base stations, and terminal devices handle local context and high-frequency lightweight tasks.\nBut the boundary is:\nElectricity is highly homogeneous; AI Tokens are highly heterogeneous.\nThe difference between electricity sources is limited, but AI Tokens from different models, tasks, and contexts have vastly different value.\nTherefore, key metrics for AI infrastructure will shift from simply \u0026ldquo;how many GPUs\u0026rdquo; to \u0026ldquo;how many AI Tokens per watt,\u0026rdquo; \u0026ldquo;cost per million AI Tokens,\u0026rdquo; \u0026ldquo;whether stable supply is available during peak hours,\u0026rdquo; and \u0026ldquo;whether edge nodes can process local context.\u0026rdquo;\n4. Like Payment Clearing Networks: Cross-Model, Cross-Supplier, Cross-Agent Settlement If a platform simultaneously connects dozens of model suppliers and serves thousands of enterprise customers, complex settlement problems arise.\nUsers see one account, one bill, one balance, but behind the scenes, multiple models, vendors, regions, and pricing systems may be invoked.\nThe platform must handle:\nUpstream model costs Downstream customer billing Department budgets Task attribution Refunds and disputes Cross-region, cross-currency, cross-supplier settlement Unified account balance Capability conversion between models This is very similar to a payment clearing network.\nBut the boundary is:\nPayment clearing networks explain cross-model conversion and unified accounts, but cannot account for model capability differences.\nMoney is highly standardized; AI Tokens are not. One million AI Tokens from a strong reasoning model and one million AI Tokens from a cheap chat model are not the same capability.\nThe Agent era will make this even more complex. In the future, a single task may be completed by multiple agents working together: one agent handles planning, one calls search, one calls code tools, one accesses the enterprise knowledge base, one calls a specialized model. They not only need to communicate but also settle the AI Tokens, tool calls, and service results each consumes.\nTherefore, cross-model settlement in the future may not just be \u0026ldquo;giving the user a bill\u0026rdquo; but evolving into some kind of AI clearing house capability: recording which model, which agent, which tool each call came from, how many AI Tokens or task results were contributed, and how costs and revenues should be shared.\nA2A, MCP, x402, machine-to-machine micropayments, stablecoin settlement, and other directions can all be seen as exploring the underlying settlement mechanisms of this Agent economy.\n5. Like CDN: Routing, Caching, and Fallback CDN delivers content to nodes closer, faster, and cheaper to users, reducing latency and cost through caching, edge access, origin protection, and fallback.\nModel capability platforms will do similar things on the scheduling side:\nSimple requests go to cheap models Complex requests go to strong models Repeated questions hit semantic cache Fallback when a model fails Dynamic routing based on latency, price, quality Model selection based on compliance requirements In the future, model routing will become a basic capability, like today\u0026rsquo;s load balancing, CDN scheduling, and database read-write separation.\nOne of the most critical technologies here is semantic caching.\nTraditional caching usually relies on exact URL, string, or hash matching; but in natural language, \u0026ldquo;what\u0026rsquo;s your return policy\u0026rdquo; and \u0026ldquo;how do I return an item\u0026rdquo; may be semantically the same question even if they look different.\nAn AI Gateway can first do fast exact matching, then use embedding and vector similarity for semantic matching: if confidence is high enough, directly return the cached answer, avoiding calling an expensive model again.\nThis brings three results: first, repeated requests no longer burn AI Tokens repeatedly; second, TTFT significantly decreases because many answers become cache reads; third, upstream model rate limit pressure drops, leaving truly expensive strong models for long-tail complex tasks.\nBut the boundary is:\nCDN explains scheduling, caching, and fallback, but AI Token scheduling also involves semantic quality, security, compliance, and responsibility.\nCDN schedules content delivery; AI Token scheduling involves generative judgment capability.\n6. Like Enterprise Governance Gateway: Permissions, Audit, and Risk Control AI Token has another easily underestimated trend: it will enter the enterprise governance system.\nEnterprises will not long allow employees to individually register model accounts, buy API keys, upload company data, connect external agents, or link business systems to untrusted models.\nThis is Shadow AI.\nWhat enterprises truly need is a unified gateway: all model calls pass through here, where identity, permissions, budgets, audit, masking, logging, cost attribution, model whitelists, and compliance policies are managed.\nSuch systems will increasingly resemble a combination of API Gateway, IAM, FinOps, security gateway, and audit systems.\nIt typically provides:\nUnified API Key management Budget management Rate limiting Cost tracking Call logs Model routing Semantic caching Fallback Prompt guard Sensitive data filtering PII/DLP detection Prompt injection defense Content safety filtering Session-level context caching Audit and compliance policies Enterprise AI Gateway is not technical pedantry, but the basic security infrastructure for organizational AI usage.\nFrom this perspective, AI Token is not just a number on a bill, but the water meter, electric meter, main switch, and audit entry point of enterprise AI governance.\n7. Like Professional Services: Tiering, Verification, and Responsibility Boundaries Finally, AI Token is also like professional services.\nThis is where the \u0026ldquo;traffic analogy\u0026rdquo; is most misleading.\nGeneral AI Tokens will become increasingly cheap. Tasks like simple Q\u0026amp;A, summarization, translation, rewriting, classification, information extraction, simple code completion, and lightweight RAG will continue to see price drops due to increased inference chips, model distillation, quantization, MoE, KV Cache optimization, prompt caching, and big-company price wars.\nBut professional AI Tokens won\u0026rsquo;t simply become commodity prices.\nIn legal, medical, financial, R\u0026amp;D, and government scenarios, what users buy is not simple inference compute, but industry corpora, professional knowledge, workflows, templates, verification mechanisms, risk warnings, audit trails, enterprise compliance, and responsibility boundaries.\nSo one sentence must be kept:\nGeneral tokens are like data; professional tokens are like expert services.\nMore bluntly:\n1GB of data is largely still 1GB of data, but between 1 million tokens and 1 million tokens, there may be the difference between an expert and a parrot.\nThis also explains why AI Tokens won\u0026rsquo;t become completely homogeneous.\nThey may become cheaper in general scenarios while commanding high multiples in professional scenarios.\n5. Aggregating the Seven Analogies: The Emergence of the Model Capacity Operator Putting the seven analogies together reveals the outline of a new role.\nIf you look only at the user side, it\u0026rsquo;s like mobile data: users need plans, credits, balances, shared packs, and overage billing.\nIf you look only at the cost side, it\u0026rsquo;s like cloud computing: platforms need cost attribution, elastic scheduling, cache optimization, budget control, and usage analytics.\nIf you look only at the infrastructure side, it\u0026rsquo;s like electricity: enterprises need stable supply, SLA, backup lines, failover, and responsibility tracking.\nIf you look only at the settlement side, it\u0026rsquo;s like a payment clearing network: multi-model, multi-supplier, multi-customer, multi-agent needs unified accounts, cross-model conversion, reconciliation, profit sharing, and dispute resolution.\nIf you look only at the scheduling side, it\u0026rsquo;s like CDN: platforms need to dynamically select models based on latency, price, quality, region, compliance, and failure status, using semantic caching and fallback to reduce cost and improve stability.\nIf you look only at the organizational side, it\u0026rsquo;s like an enterprise governance gateway: enterprises need permissions, budgets, audit, masking, logs, model whitelists, and compliance policies.\nIf you look only at the professional side, it\u0026rsquo;s like professional services: users need not just cheap tokens, but verifiable, deliverable, accountable professional results.\nSo the seven analogies ultimately point not to an abstract concept, but to a very concrete set of operational capabilities:\nPlan design capability\nToken-native metrics management (TTFT/TPS/Cost per Token/Tokens per Watt) Cost governance capability Stable supply capability AI Grid / edge inference organization capability Cross-model settlement capability Inter-agent clearing capability Intelligent scheduling capability Semantic caching capability Enterprise governance capability Compliance audit capability Professional verification capability Result accountability capability\n= Model Capacity Operator This is what I call the Model Capacity Operator.\nIt\u0026rsquo;s not a simple API forwarder.\nIt doesn\u0026rsquo;t necessarily own models.\nBut it operates the gateway, metering, routing, ledger, quality, compliance, and responsibility boundaries of model capacity.\nThe following diagram breaks down the capability structure of the \u0026ldquo;Model Capacity Operator\u0026rdquo;: from infrastructure, to operational governance, to value delivery—the real moat is not how many models you can access, but whether you can operate model capacity into stable, trustworthy, and deliverable results.\nEarly model relay stations competed on model count, price, and API compatibility. Whoever could offer more models, cheaper prices, and OpenAI API compatibility would attract developers more easily.\nBut the next stage will be about operational capability.\nThat is, the real question becomes:\nWho has the ability to operate model calls—from individual API requests—into stable, trustworthy, deliverable, auditable, billable, and accountable model capacity?\nThis is the core reason why Model Capacity Operators will emerge.\n6. Capability Boundaries, Business Models, and Why Model Capacity Operators Emerge After introducing the concept of the Model Capacity Operator, three questions remain: What can it do? What can\u0026rsquo;t it do? How does it make money?\n1. Capability Boundaries: Can Operate Capacity, But Cannot Eliminate Responsibility Model Capacity Operators can do a lot.\nThey can provide a unified API, consolidating multiple models into a single gateway.\nThey can build an LLM Router that selects models based on task type, cost, latency, context length, data sensitivity, quality requirements, and compliance region.\nThey can implement semantic caching, so repeated or similar questions no longer burn AI Tokens repeatedly.\nThey can set up fallback, automatically switching to backup models when the primary model fails, hits rate limits, or degrades in quality.\nThey can perform cost attribution, allocating AI Token consumption to departments, projects, customers, orders, and tasks.\nThey can implement budget control, letting enterprises know who is spending money, where it\u0026rsquo;s going, and whether they\u0026rsquo;re over budget.\nThey can provide model whitelists, data masking, content safety, prompt injection defense, PII/DLP detection, RBAC, log retention, compliance auditing, and content watermarking.\nThey can also handle cross-model settlement, inter-agent settlement, refund dispute resolution, and supplier reconciliation.\nIt\u0026rsquo;s important to distinguish between model providers and Model Capacity Operators. Model providers are mainly responsible for foundational model training, parameter iteration, and base model capability. Model Capacity Operators are closer to system-layer deployment, governance, and stewardship—they integrate models into enterprise workflows and provide access control, audit logs, DLP, content watermarking, call records, and responsibility attribution. As regulations in different regions increasingly differentiate between roles like provider, deployer, and distributor, this boundary will become even more important.\nBut Model Capacity Operators also have limits.\nThey cannot guarantee that all model outputs are absolutely correct.\nThey cannot evade all responsibility by claiming \u0026ldquo;I\u0026rsquo;m just a relay.\u0026rdquo;\nThey cannot package opaque model sources as a low-price advantage.\nThey cannot treat high-risk professional tasks as simple general chat.\nThey cannot hide enterprise data security, compliance auditing, and content responsibility inside a black box.\nMore importantly, they must acknowledge:\nBehind every AI Token is a judgment, and judgments carry quality, risk, and responsibility.\nSo the capability boundary of a Model Capacity Operator is not \u0026ldquo;how many models I can access,\u0026rdquo; but \u0026ldquo;to what extent I can make model capacity controllable, auditable, billable, deliverable, and accountable.\u0026rdquo;\n2. Possible Business Models Model Capacity Operators may have multiple business models in the future.\nFirst, AI Token packs. This is closest to today\u0026rsquo;s relay model: users purchase credits and are deducted based on consumption. The difference is that platforms that survive long-term must have upstream authorization, genuine models, transparent billing, and data security.\nSecond, enterprise model gateway subscription. Enterprises purchase AI Gateway capabilities on a monthly or yearly basis, including unified gateway, permissions, budgets, logs, audit, model whitelists, masking, routing, and compliance policies.\nThird, task-based pricing. Users no longer buy AI Tokens directly, but purchase task results—such as document processing, customer service tickets, code review, contract review, knowledge base Q\u0026amp;A, and automated agent execution.\nFourth, industry task packs. Targeting legal, medical, financial, education, and government scenarios, packaging model calls, industry knowledge, templates, verification mechanisms, and audit trails into professional services.\nFifth, private deployment, sovereign cloud, and hybrid cloud services. Enterprises place sensitive tasks on local, private, or sovereign cloud infrastructure, while complex low-frequency tasks go to cloud-based strong models. The platform handles unified scheduling and cost governance. For multinational enterprises, finance, government, and healthcare scenarios, data residency, access control, audit evidence, and supply chain explainability themselves become reasons to pay.\nSixth, cost optimization revenue sharing. Platforms help enterprises reduce AI Token costs through semantic caching, model routing, batch processing, context compression, and local model substitution, then share in the savings.\nSeventh, compliance audit and content watermarking services. When regulations require traceability of model sources, generated content, data processing, and call chains, audit capability itself becomes a product.\nEighth, Agent invocation and clearing services. When multiple agents, models, and tools work together on a task, the platform can provide call records, cost allocation, revenue settlement, dispute resolution, and responsibility tracking.\nNinth, multi-layer platform and professional services. Mature platforms won\u0026rsquo;t rely solely on API price arbitrage. They may form a combination of IaaS, PaaS, SaaS, and professional services: the bottom layer provides GPU or inference resources; the middle layer provides LLM interfaces, fine-tuning environments, model gateways, and development platforms; the top layer provides white-label enterprise Copilot, no-code AI building tools, industry agents, and algorithm marketplaces. Revenue will also expand from simple pay-per-use to subscriptions, revenue sharing, private deployment, custom development, and compliance consulting.\n3. Why This Role Must Emerge The Model Capacity Operator is not a concept pulled from thin air, but the result of multiple converging pressures.\nFirst, the number of models keeps growing, and users cannot manage all of them themselves.\nSecond, model pricing is becoming increasingly complex, and users cannot study price tables every day.\nThird, enterprise AI usage is becoming ubiquitous, and organizations must govern Shadow AI.\nFourth, AI Token costs will shift from technical costs to financial costs—enterprises need budgeting, attribution, and optimization.\nFifth, agent call chains will grow longer, with a single task potentially spanning multiple models, tools, knowledge bases, and external services.\nSixth, regulation will increasingly focus on data, content, model sources, and call chains.\nSeventh, what users ultimately want to buy is task results, not underlying AI Tokens.\nTherefore, in the long run, the model call market will not stay at \u0026ldquo;who sells cheaper.\u0026rdquo;\nIt will move toward \u0026ldquo;who can operate model capacity more stably, more compliantly, at lower cost, with better scheduling, clearer settlement, and greater accountability.\u0026rdquo;\n7. Development Trends of Model Capacity Operations Following this line of reasoning, I believe model capacity operations will exhibit nine trends.\n1. Gray-Market Relays Recede, Compliant Aggregation Rises Platforms relying on gray key pools, unauthorized resale, fake models, opaque proxies, no data agreement, and no audit capabilities will find it increasingly difficult to survive.\nBut authorized API aggregation platforms, cloud vendor model platforms, enterprise AI Gateways, and industry capability platforms will become increasingly important.\nIn short:\nGray-market relays will be compressed, but model aggregation will not disappear.\n2. General AI Token Unit Price Drops, But Total Consumption Rises General AI Tokens will continue to get cheaper.\nReasons include increased inference chips, model distillation, quantization, MoE, KV Cache optimization, open-source model catch-up, prompt caching, and big-company price wars.\nBut total consumption won\u0026rsquo;t necessarily decline.\nBecause agents will automatically break down tasks and call models in multiple rounds; long contexts will increase input; RAG will stuff in large amounts of background material; multimodality will process images, audio, and video; tool calls will add steps; multi-agent collaboration will amplify call volume.\nSo the future may present a seemingly contradictory phenomenon:\nEach token is cheaper, but each task consumes more tokens behind the scenes.\n3. Pricing Moves from Pay-Per-Use to Plans and Task Packaging Today, everyone is still comparing prices per million AI Tokens.\nIn the future, it\u0026rsquo;s more likely to become: AI office plans, code assistant plans, agent call packs, enterprise knowledge base plans, document processing packs, after-sales diagnostic task packs, and industry model service packs.\nWhat users ultimately care about is not token unit price, but:\nHow much does task completion cost? Can it be completed reliably? Is the quality dependable? Is it auditable? Who is responsible if something goes wrong? This will transform AI Tokens from \u0026ldquo;pay-per-use resources\u0026rdquo; into \u0026ldquo;capability packages.\u0026rdquo;\n4. Model Routing Becomes a Foundational Capability Future AI applications will not be tied to a single model, but connected to a model capability pool.\nSystems will automatically select models based on cost, latency, accuracy, context length, multimodal capability, data sensitivity, compliance region, and fallback requirements.\nSimple tasks use cheap models, complex tasks use strong models; public data uses public models, sensitive data uses private models; when the primary model fails, automatically switch to a backup.\nModel routing will become as foundational as today\u0026rsquo;s load balancing, CDN scheduling, and database read-write separation.\n5. Enterprise AI Gateway Becomes the Organizational AI Access Point Enterprises will increasingly care about the AI call gateway.\nIt needs to centrally manage identity, permissions, budgets, model access, logs, audit, sensitive data, cost attribution, fallback, and compliance policies.\nWithout this gateway, a large amount of Shadow AI will emerge within organizations: employees buying their own keys, uploading their own data, connecting external tools, and sending business data to unknown models.\nThis is unmanageable for enterprises.\n6. Professional Vertical AI Tokens Will Command Premium Multiples General AI Tokens will drop in price, but professional vertical models won\u0026rsquo;t simply compete on low price.\nBecause professional AI Tokens sell more than just inference compute—they include industry corpora, professional knowledge, workflows, templates, verification mechanisms, risk warnings, audit trails, enterprise compliance, and responsibility boundaries.\nOrdinary tokens sell compute.\nProfessional tokens sell experience, process, accountability, and verification.\n7. AI-Native Metrics Become Service Quality Standards Traditional networks look at ping, bandwidth, and packet loss. Traditional API gateways mostly promise uptime. AI services will increasingly look at TTFT, TPS, context capacity, cache hit rate, Cost per Token, Tokens per Watt, Goodput, and task success rate.\nThis means that future Model Capacity Operators won\u0026rsquo;t just sell interfaces—they will need to publish, monitor, and optimize a full set of service quality metrics, just like cloud providers and telecom carriers.\nWhoever can get users their first token faster, generate long answers more stably, complete tasks at lower cost, and deliver effective output with sufficient quality within agreed latency SLOs will be more competitive. Future SLAs won\u0026rsquo;t just say \u0026ldquo;system available\u0026rdquo;—they will increasingly mean \u0026ldquo;model capability available, response available, result available.\u0026rdquo;\n8. The Agent Economy Will Drive the Emergence of an AI Clearing Layer As AI applications move from single-turn chat to multi-agent collaboration, settlement problems will be amplified.\nA complex task may invoke multiple models, multiple tools, multiple external agents, and multiple knowledge bases. The end user sees only one result, but behind the scenes, the platform must record each participant\u0026rsquo;s AI Token consumption, tool contribution, result quality, and revenue distribution.\nSo in the future, Model Capacity Operators may take on a role similar to an AI clearing house: managing both calls and ledgers, both model routing and cross-agent cost allocation, settlement, and dispute resolution.\n9. The Moat Shifts from \u0026ldquo;Number of Models\u0026rdquo; to \u0026ldquo;Governance Capability\u0026rdquo; Early platforms liked to boast about how many models they had integrated.\nBut in the future, what really matters may not be the count, but:\nCan you prove the model source is genuine? Can you provide a stable SLA? Can you optimize costs? Can you do semantic caching? Can you automatically fallback? Can you satisfy enterprise audit requirements? Can you implement content watermarking? Can you handle cross-model settlement? Can you bear responsibility boundaries? More models mean harder management; more calls mean more important governance; more specialized scenarios mean heavier responsibility.\nThis is the long-term moat of the Model Capacity Operator.\n8. How Should Different Roles Respond? If this trend holds, different roles should adopt different strategies.\nThis diagram serves as an action map for this section: gray-market relays recede, compliant aggregation strengthens, Model Capacity Operators emerge—different roles have different priorities.\n1. Individual Users: You Can Use Relays, But Don\u0026rsquo;t Blindly Trust Individual users can use relay services in small amounts, but don\u0026rsquo;t trust blindly.\nDon\u0026rsquo;t make large deposits. Don\u0026rsquo;t upload sensitive materials. Don\u0026rsquo;t upload company code, contracts, or customer information. Don\u0026rsquo;t tie critical tasks to a single small platform.\nIt\u0026rsquo;s best to keep an official API or second supplier as backup, and pay attention to whether the platform discloses upstream sources, pricing, privacy policies, and SLA.\nIn short:\nIndividual users can treat relay stations as tools, but not as infrastructure.\n2. Developers and Small Teams: Build Model Abstraction Early Developers and small teams should implement provider abstraction as early as possible, and not hardcode business logic to a single model.\nCode should support multi-model switching. Critical tasks should have fallback. Simple tasks should use cheap models. Complex tasks should use strong models. Sensitive tasks should use trusted cloud or local models. Also track costs and call logs.\nThe future belongs not to those who bind most deeply to one model, but to those who are best at switching, routing, and managing models.\n3. Enterprises: Build AI Gateway Thinking as Soon as Possible Enterprises need to avoid Shadow AI.\nEmployees should not be allowed to privately register model accounts, purchase API keys, upload corporate data, connect external agents, or link business systems to untrusted models.\nEnterprises should establish a unified AI Gateway to manage model access, account permissions, department budgets, sensitive data, call logs, audit records, model whitelists, data cross-border transfers, cost attribution, and compliance policies.\nEnterprise AI Gateway is not technical pedantry, but the basic security infrastructure for organizational AI usage in the future.\n4. Entrepreneurs: Don\u0026rsquo;t Just Do Low-Price Relays—Do Capability Operations Simply reselling tokens has no long-term moat.\nUpstream can tighten authorization, big companies can cut prices, cloud vendors can aggregate, user switching costs are low, and gray models are unsustainable.\nEntrepreneurial opportunities are more likely in model routing, cost optimization, enterprise gateways, industry knowledge bases, task-based plans, private deployment, vertical industry agents, audit and compliance tools, and professional workflows.\nIn the future, the most profitable business won\u0026rsquo;t be selling tokens, but packaging tokens into reliable results.\n5. Vertical Industry Practitioners: Opportunity Lies in Professional Scenarios, Not General Low Prices Vertical industry practitioners should spend less time on \u0026ldquo;which model is cheapest\u0026rdquo; and more time on \u0026ldquo;which industry tasks can be reliably completed by model capability.\u0026rdquo;\nIf you can combine industry knowledge, process templates, enterprise data, audit trails, result verification, and responsibility boundaries, there is an opportunity to transform token consumption into professional capability services.\nThis is not a low-price model war—it is the productization of industry expertise.\n9. Conclusion: Relay Is Just the Starting Point, Model Capacity Operations Are the Long-Term Direction The explosion of model proxies, API relays, and token agents is not an isolated gray-market phenomenon, but an early signal that AI Tokens are moving toward model capacity operations.\nToday\u0026rsquo;s chaos tells us two things.\nFirst, demand already exists. Users genuinely need unified gateways, unified billing, multi-model access, cost management, and lower usage barriers.\nSecond, early-stage supply is still rough. Gray key pools, unauthorized resale, fake models, black-box pipelines, data risks, and lack of audit capabilities all indicate that this market has not yet matured into a legitimate form.\nBut in the long run, gray-market relays will be cleaned up, while model aggregation will not disappear; low-price key pools will vanish, while unified gateways, unified billing, model routing, enterprise audit, content watermarking, and industry plans will become increasingly important.\nAI Tokens will become plan-based like mobile data, cost-optimized like cloud resources, pursuit of stable supply like electricity, requiring reconciliation and settlement like payment clearing networks, needing scheduling and caching like CDN, entering permission, budget, and audit systems like enterprise governance gateways, and forming responsibility boundaries and price premiums in high-value scenarios like professional services.\nBut it will not become as homogeneous as these resources.\nBecause data transfers information, while Tokens generate judgments.\nSo the future of AI Tokens is not simply about price drops, nor about relay stations disappearing—it is about moving from gray-market relays to compliant operations, from single-model calls to multi-model scheduling, from token metering to AI capability plans, from API forwarding to model capacity operations.\nWhat will truly be valuable in the future is not whose tokens are cheaper, nor who can relay more models, but who can operate the model capability behind those tokens into stable, trustworthy, deliverable, auditable, billable, and accountable results.\n","permalink":"/en/posts/ai-token-trend-prediction/","summary":"\u003cblockquote\u003e\n\u003cp\u003e\u003cstrong\u003eTL;DR\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis article is not about \u0026ldquo;which API is cheaper,\u0026rdquo; but about \u003cstrong\u003eAI Tokens\u003c/strong\u003e becoming the settlement unit for model capacity—and who will manage the gateway, routing, ledger, quality, and responsibility.\nCore judgment: \u003cstrong\u003eModel reselling is not the endgame; model capacity operators are the long-term direction.\u003c/strong\u003e\nGray-market reselling will recede, but unified APIs, multi-model routing, semantic caching, compliance auditing, enterprise governance, and cost optimization will remain. The future competition will not be about \u0026ldquo;how many models you can access,\u0026rdquo; but whether you can operate model capacity into stable, trustworthy, deliverable, auditable, billable, and accountable results.\u003c/p\u003e","title":"From Model Reseller to Model Capacity Operator: Predicting the Development Trend of AI Tokens"}]