AgentOS · TypeScript AI Agent Framework
Agents that remember, forge their own tools, and survive long-running sessions. Persistent cognitive memory, optional HEXACO personality, multi-agent orchestration, and one dispatch interface across 13 LLM providers. Apache-2.0.
Benchmarks * Website * Docs * npm * Discord * Blog
AgentOS is an open-source TypeScript framework for AI agents that remember, adapt, and write their own tools.
- Top open-source memory benchmarks: 85.6% on LongMemEval-S at $0.0090/correct (gpt-4o), and 70.2% on LongMemEval-M, the only open-source library above 65% on M with reproducible methodology.
- Runtime tool forging. An agent writes a JavaScript function with JSON Schemas for its input and output, the declared test cases run it, an LLM judge reviews the result, and on approval the tool joins the session's tool list. Forged code runs in an in-process
node:vmcontext or, withQuickJSExecutor, in a QuickJS WebAssembly instance of its own for each call. - Persistent cognitive memory with Ebbinghaus decay and 8 neuroscience-backed mechanisms, among them retrieval-induced forgetting, reconsolidation and source-confidence decay.
- Optional HEXACO personality, 6 orchestration strategies, guardrails, and voice across 13 LLM providers; 100+ extensions and 88 skills ship as separate packages with registries that load them.

Runtime tool forging + multi-agent collaboration. Reproduce with node examples/emergent-hierarchical-spawning.mjs.
Install
npm install @framers/agentos
import { agent } from '@framers/agentos';
const tutor = agent({
provider: 'anthropic', // resolves to claude-sonnet-4-6 (provider default)
// model: 'claude-opus-4-8', // pin a specific model to override the default
instructions: 'You are a patient CS tutor.',
personality: { openness: 0.9, conscientiousness: 0.95 },
});
// Provider auto-detected from env when `provider` is omitted.
// Cognitive memory, sentiment tracking and metaprompts: runtime: 'gmi' (see GMIs below).
const session = tutor.session('student-1');
await session.send('Explain recursion with an analogy.');
await session.send('Can you expand on that?'); // remembers context
Full quickstart * Examples cookbook * API reference
Sessions. A session keeps the whole conversation whether memory is on or off: each send() records its tool calls, their results and the model's signed thinking, and every later request replays them. stream() records its turn as the prompt and the final text (with runtime: 'gmi', every step, as send() does). History is capped at about 120K tokens by default. Set history: false for a stateless session, set history: { maxTokens } to change the cap, and call reseed() to replace the history with a shorter set of messages you build yourself. close() ends a session and frees its history: the next session(id) with that id starts empty, and agent.usage(id) still reports what the id spent.
const stateless = agent({ model, memory: false, history: false }).session('job-1');
const bounded = agent({ model, history: { maxTokens: 60_000 } }).session('job-2');
bounded.reseed([{ role: 'user', content: 'compact resume snapshot' }]);
Emergent Design
Three things accumulate across a session and compose into behavior: memory (what was said, decided, retrieved), the tool surface (which grows when an agent forges a tool the judge approves), and an optional HEXACO personality vector that shapes the prompt and scales memory encoding. Each is configurable and observable.
Runtime tool forging. When no tool covers a sub-task, the agent calls forge_tool with a chain of existing tools or a JavaScript function, JSON Schemas for input and output, and test cases. Code mode is off until the host sets emergentConfig.allowSandboxTools. Source that uses eval, require or process is rejected before it runs. Forged code runs in an in-process node:vm context by default, with a 5 s timeout, and node:vm is not a security boundary; with QuickJSExecutor, each call runs in a QuickJS WebAssembly instance of its own with its memory limited. A separate LLM judge reviews the candidate with its test results, and an approved tool joins the session's tool list. A forged tool exports as a SKILL.md skill. Emergent capabilities ->
HEXACO personality (optional). Off by default. agent() writes the trait values into the system prompt as directives, one for each trait above 0.65 or below 0.35. A cognitive memory manager given the same traits scales its encoding and seven of its eight mechanisms by them, and an AgentGraph branches on a trait with addPersonalityEdge(). HEXACO docs ->
Soul files. Identity, voice, hard limits, and HEXACO scores can live in a SOUL.md workspace. Its memory/ directory is a markdown wiki (an index.md catalog plus entities/, concepts/, log/ pages with [[wikilinks]]) that is the agent's long-term memory: markdown is the source of truth, the vector/graph index is rebuilt from it, and souledAgent() wires it end to end. Soul Files ->
import { souledAgent } from '@framers/agentos';
const aria = await souledAgent({ provider: 'anthropic', soul: '~/.agentos/agents/aria' });
Generalized Mind Instances (GMIs)
On the full runtime, every session is served by a GMI: a persistent agent with its own persona, mood, conversation history and reasoning trace; with agent({ runtime: 'gmi' }) an agent's sessions are GMIs too. agent() without runtime: 'gmi' is the lightweight helper; it calls the model with a prompt and keeps session history. A GMI of the full runtime runs a turn loop around the same model, tools, guardrails and cognitive memory:
- Sentiment → metaprompts. When a persona enables sentiment tracking, every user turn is scored; sustained frustration or confusion fires recovery metaprompts, and a self-reflection metaprompt re-reads the GMI's mood and task context from evidence.
- Mood-weighted memory. With cognitive memory attached, each exchange is encoded with the GMI's current mood and recalled with emotional congruence in the score.
- Self-modification tools. With
selfImprovement.enabled, the runtime registersadapt_personality,manage_skills,create_workflowandself_evaluate;adapt_personalitychanges the running GMI's traits within bounds, and a mutation store records the changes when a storage adapter andpersistWithDecayare configured. - A reasoning trace of the last 500 decisions by default (
reasoningTraceConfigon the persona or the runtime's default), and persona overlays per session.
import { AgentOS, AgentOSResponseChunkType, BUILT_IN_PERSONAS } from '@framers/agentos';
// AgentOS.create() reads persona files from ./personas by default. Personas can
// also be given inline, as parsed JSON or code-built objects; here, the five the
// package ships. A custom loader covers any other source.
const agentos = await AgentOS.create({ personas: BUILT_IN_PERSONAS });
for await (const chunk of agentos.processRequest({
userId: 'user-42', sessionId: 'research-q1', selectedPersonaId: 'v_researcher',
textInput: 'Summarize the open incidents from this week.',
})) {
if (chunk.type === AgentOSResponseChunkType.TEXT_DELTA) process.stdout.write(chunk.textDelta);
}
agent({ runtime: 'gmi' }), also exported as gmi(), builds its GMIs in process from the agent's options, with no AgentOS runtime. Of the parts above, these GMIs run the reasoning trace, the mood-weighted memory when memory is on, and sentiment tracking with the five preset event metaprompts (frustration recovery, confusion clarification, satisfaction reinforcement, error recovery and engagement boost) when their profile turns them on. The 'full' profile turns on cognitive memory, sentiment tracking and those five metaprompts; 'light', the default, keeps the reasoning trace and adds memory only when memory is set. The self-reflection metaprompt, the self-modification tools, persona overlays and the runtime's guardrails, retrieval and channels do not run on this path. This tutor's GMIs run every part this path has:
import { agent } from '@framers/agentos';
const tutor = agent({
runtime: 'gmi',
cognition: 'full',
provider: 'openai',
instructions: 'You are a patient CS tutor.',
memory: { embedding: { provider: 'openai' } }, // cognitive memory embeds with text-embedding-3-small
});
const session = tutor.session('student-1', { userId: 'student-7f3a' }); // memory is scoped to the user id
await session.send('Explain recursion with an analogy.');
What a GMI adds over a plain agent →
Memory Benchmarks
gpt-4o reader, gpt-4o-2024-08-06 judge, full N=500, single-CLI reproduction with bootstrap 95% CIs and per-benchmark judge-FPR probes.
- LongMemEval-S: 85.6% at $0.0090/correct, 3,558 ms p50: +1.4 points over Mastra OM gpt-4o (84.23%), 0.4 behind Emergence.ai's closed-source 86%. The highest publicly reproducible open-source number at
gpt-4o. - LongMemEval-M: 70.2% (1.5M-token haystacks, 500 sessions): the only open-source library above 65% on M with reproducible methodology.
Full leaderboard -> * Transparency audit -> * LongMemEval paper (Wu et al., ICLR 2025)
Why AgentOS
| vs. | AgentOS differentiator |
|---|---|
| LangChain / LangGraph | Cognitive memory (8 neuroscience-backed mechanisms), HEXACO personality, runtime tool forging |
| Vercel AI SDK | Multi-agent teams (6 strategies), 7 vector backends, guardrails, voice/telephony, zero-config prompt caching |
| CrewAI / Mastra | Unified orchestration (DAGs + graphs + missions), personality-driven routing, published reproducible numbers on LongMemEval-S (85.6%) and LongMemEval-M (70.2%) with full methodology disclosure |
Key Features
| Category | Highlights |
|---|---|
| LLM Providers | 13 (11 by API key or base URL + 2 local CLI): OpenAI, Anthropic, Gemini, Groq, Ollama, OpenRouter, Requesty, LiteLLM, Together, Mistral, xAI, Claude CLI, Gemini CLI. Plus image/video/audio generation providers. |
| Prompt Caching | Zero config on every provider: automatic Anthropic breakpoints incl. multi-turn history (direct + OpenRouter) * OpenAI cache-key routing * normalized cache usage + leak detection * per-call TTL/opt-out * guide |
| Cognitive Memory | 8 mechanisms: reconsolidation, retrieval-induced forgetting, involuntary recall, FOK, gist extraction, schema encoding, source decay, emotion regulation |
| HEXACO Personality | 6 traits modulate memory, retrieval bias, response style |
| GMI Runtime | Per-session persona, mood and reasoning trace * sentiment-triggered metaprompts (opt-in per persona) * mood-weighted memory bridge * bounded adapt_personality trait changes (opt-in) |
| RAG Pipeline | 7 vector backends * 4 retrieval strategies * GraphRAG * HyDE * Cohere rerank-v3.5 |
| Multi-Agent Teams | 6 coordination strategies * manager delegation and specialist spawning * panel quorum on the parallel strategy * HITL approval gates |
| Orchestration | workflow() DAGs * AgentGraph router, personality and discovery edges * mission() graphs from plan templates * checkpointing |
| Guardrails | 5 security tiers * 6 packs (PII, ML classifiers, topicality, code safety, grounding, content policy) |
| Emergent Capabilities | Runtime tool forging * 4 self-improvement tools * tiered promotion * skill export |
| Voice & Telephony | ElevenLabs, Deepgram, Whisper * Twilio, Telnyx, Plivo |
| Channels | 37 platform adapters (Telegram, Discord, Slack, WhatsApp, webchat, ...) |
| Observability | OpenTelemetry * usage ledger * cost guard * circuit breaker |
Multi-Agent in 6 Lines
import { agency } from '@framers/agentos';
const team = agency({
strategy: 'graph',
agents: {
researcher: { provider: 'anthropic', instructions: 'Find relevant facts.' }, // -> claude-sonnet-4-6
writer: { provider: 'openai', instructions: 'Summarize clearly.', dependsOn: ['researcher'] }, // -> gpt-4o
reviewer: { provider: 'gemini', instructions: 'Check accuracy.', dependsOn: ['writer'] }, // -> gemini-2.5-flash
},
});
const result = await team.generate('Compare TCP vs UDP for game networking.');
Strategies: sequential, parallel, debate, review-loop, hierarchical, graph. With hierarchical + emergent: { enabled: true }, the manager forges new sub-agents at runtime. Every roster agent can set its own provider, model, apiKey and effort; a parallel agency can require a provider quorum (quorum: { minProviders: 2 }) before it synthesizes. Multi-agent docs ->
Ecosystem
| Package | Role |
|---|---|
@framers/agentos |
Core runtime: agents, cognitive memory, orchestration, guardrails, voice, 13 LLM providers. Apache-2.0. |
@framers/agentos-extensions |
100+ first-party extensions: channel adapters, tool packs, integrations, guardrail packs. |
@framers/agentos-extensions-registry |
Discovery + auto-loader for the extensions catalog. |
@framers/agentos-skills |
88 curated SKILL.md skills. |
@framers/agentos-skills-registry |
Discovery + auto-loader for skills; where promoted forged tools land. |
@framers/agentos-bench |
Open benchmark harness: bootstrap 95% CIs, judge-FPR probes, per-case run JSONs. MIT. |
@framers/sql-storage-adapter |
Cross-platform SQL persistence: SQLite, Postgres, IndexedDB, Capacitor SQLite. |
paracosm |
AI agent swarm simulation on AgentOS. Live demo. |
wunderland |
Batteries-included CLI + daemon over the AgentOS registries (preview). Apache-2.0. |
Extensions load from the manifest a host passes to AgentOS.create({ extensionManifest }); createCuratedManifest() from @framers/agentos-extensions-registry builds one from the curated extensions that are installed, and a runtime without a manifest loads none. Extensions architecture ->
Configure API Keys
Three layers, highest priority first: inline apiKey on the call, a module-level setDefaultProvider() at boot, or environment-variable auto-detection (OPENAI_API_KEY, ANTHROPIC_API_KEY, and the rest, resolved in priority order and reorderable with setProviderPriority([...])). A comma-separated list of keys rotates per request on the providers that support it (Key rotation).
Full credential resolution + default models per provider ->
API Surfaces
agent(): lightweight stateful agent. Prompts, sessions, personality, hooks, tools, memory throughmemoryProviderhooks. Withruntime: 'gmi'(orgmi()), every session is a GMI built from the same options, with cognitive memory, sentiment tracking and metaprompts per itscognitionprofile.agency(): multi-agent teams built fromagent()members, with HITL approval gates, run limits, structured output, provenance and, on the hierarchical strategy, specialists spawned at runtime. Itsguardrailsandragoptions are reported or logged and not applied, andvoice.enabledserves the agency as JSON text over a local WebSocket. It wires no channels; channel adapters run on the full runtime or withChannelRouter.generateText()/streamText()/generateObject()/generateImage()/generateVideo()/generateMusic()/performOCR()/embedText(): low-level multi-modal helpers with native tool calling.workflow()/AgentGraph/mission(): three orchestration authoring APIs over one graph runtime.
Provider fallback is on by default for generateText(), streamText(), agent() and agency(): when a call fails with a retryable error, it is retried on the other providers whose keys are in the environment. Pass fallbackProviders: [] to turn it off, or a list to set the chain yourself. The GMIs of the full runtime (processRequest()) call their provider without a fallback chain. The GMIs of agent({ runtime: 'gmi' }) call it through a completion gateway built from the agent's fallbackProviders, which moves to the next provider when a call fails with a retryable error before its first output, and not after it.
Full API reference -> * High-Level API guide ->
Documentation & Community
- Benchmarks: benchmark tables, 95% confidence intervals, methodology audit
- Architecture: system design, layer breakdown
- Cognitive Memory: 8 mechanisms with 30+ APA citations
- RAG Configuration: vector stores, embeddings, sources
- Guardrails: 5 tiers, 6 packs
- Voice Pipeline: TTS, STT, telephony
- Blog: engineering posts, benchmark publications, transparency audits
- Discord * GitHub Issues * Wilds.ai (AI game worlds powered by AgentOS)
Contributing
git clone https://github.com/framerslab/agentos.git && cd agentos
pnpm install && pnpm build && pnpm test
We use Conventional Commits. Project guides:
| Guide | What |
|---|---|
| Contributing | Development setup, commit and pull request rules, review threads, contribution licensing |
| Adding an LLM provider | Provider interface, acceptance checklist, sponsorship and disclosure |
| Release guide | How a merge to master becomes an npm release |
| Agent instructions | Commands and conventions for coding agents |
| Maintainers | Who reviews and merges changes |
| Code of Conduct | Community standards |
| Security Policy | Reporting vulnerabilities privately |
| Support | Where to get help |
| Sponsors | Funding, sponsor placement and disclosure |
Startups & Partnerships
AgentOS is Apache-2.0 and free. We integrate any quality provider on technical merit, and partners and sponsors are featured in the README and docs, labeled as such. Companies engage through partner startup programs, sponsorship, or a provider integration. See SPONSORS.md.
Programs & partners
| Partner | Type | Provides | Since |
|---|---|---|---|
| Startup Program | Speech-to-text + text-to-speech credits, go-to-market | 2026 |
Ways to engage
| Track | What it is | Where |
|---|---|---|
| Sponsor | Fund development. Disclosed logo placement + release-notes credit. | SPONSORS.md |
| Provider integration | Ship your model or API as a supported provider. Free, on technical merit. | Provider guide |
Interested? Email team@frame.dev.

