🦌 DeerFlow - 2.0
English | 中文 | 日本語 | Français | Русский | Português
On February 28th, 2026, DeerFlow claimed the 🏆 #1 spot on GitHub Trending following the launch of version 2. Thanks a million to our incredible community — you made this happen! 💪🔥
DeerFlow (Deep Exploration and Efficient Research Flow) is an open-source super agent harness that orchestrates sub-agents, memory, and sandboxes to do almost anything — powered by extensible skills.
https://github.com/user-attachments/assets/a8bcadc4-e040-4cf2-8fda-dd768b999c18
[!NOTE] DeerFlow 2.0 is a ground-up rewrite. It shares no code with v1. If you're looking for the original Deep Research framework, it's maintained on the
1.xbranch — contributions there are still welcome. Active development has moved to 2.0.
Official Website
Learn more and see real demos on our official website. The landing-page case studies open as allowlisted, read-only showcases without requiring a sign-in.
Sister Projects
- LLM Space - Meet our secret weapon behind DeerFlow — one desktop tool to prototype agent ideas, inspect each harness step, replay failures, and benchmark performance.
Coding Plan from ByteDance Volcengine
- We strongly recommend using Doubao-Seed-2.0-Code, DeepSeek v3.2 and Kimi 2.5 to run DeerFlow
- Learn more
- 中国大陆地区的开发者请点击这里
InfoQuest
InfoQuest reader, web search, and image search use a 30-second HTTP connect/read
inactivity timeout. The crawl timeout and navigation_timeout settings remain
separate server-side options; they do not control the local HTTP timeout.
DeerFlow has newly integrated the intelligent search and crawling toolset independently developed by BytePlus — InfoQuest (supports free online experience)
Table of Contents
- 🦌 DeerFlow - 2.0
- Official Website
- Sister Projects
- Coding Plan from ByteDance Volcengine
- InfoQuest
- Table of Contents
- One-Line Agent Setup
- Quick Start
- From Deep Research to Super Agent Harness
- Core Features
- Recommended Models
- Embedded Python Client
- Projects
- Scheduled Tasks
- Terminal Workbench (TUI)
- Documentation
- ⚠️ Security Notice
- Contributing
- License
- Acknowledgments
- Star History
One-Line Agent Setup
If you use Claude Code, Codex, Cursor, Windsurf, or another coding agent, you can hand it the setup instructions in one sentence:
Help me clone DeerFlow if needed, then bootstrap it for local development by following https://raw.githubusercontent.com/bytedance/deer-flow/main/Install.md
That prompt is intended for coding agents. It tells the agent to clone the repo if needed, choose Docker when available, and stop with the exact next command plus any missing config the user still needs to provide.
Quick Start
Configuration
Operators can extend lead-agent, subagent, and DeerMem extraction prompts with literal prepend/append configuration without editing source templates. See prompt overlays.
Optional per-model request_admission
paces requests to help stay within provider request-per-minute limits.
It is disabled by default; see the linked guide to enable it.
Enabling it disables exposed SDK retries and the Claude and Codex adapters'
internal retry loops so middleware retries pass through request admission again.
Middleware retries also cover HTTP 529 overload responses.
A warning identifies retry_max_attempts values overridden by admission.
For Google's official Gemini OpenAI-compatible endpoint, use the Gemini reasoning profile.
For MindIE XML tool calls, see the argument parsing and newline compatibility guide. Both synchronous and asynchronous streams retain this compatibility: tool-enabled streams simulate chunks from a non-streaming response, while no-tool streams stay native.
-
Clone the DeerFlow repository
git clone https://github.com/bytedance/deer-flow.git cd deer-flow -
Run the setup wizard
From the project root directory (
deer-flow/), run:make setupThis launches an interactive wizard that guides you through choosing an LLM provider, optional web search, and execution/safety preferences such as sandbox mode, bash access, and file-write tools. It generates a minimal
config.yamland writes your keys to.env. Takes about 2 minutes.The wizard also lets you configure an optional web search provider, or skip it for now.
Brave web search preserves valid entries in mixed result lists. Malformed response containers or lists containing no usable entries return a structured format error; missing, null, or empty results keep the existing "No results found" response. Format errors also log the malformed container's path and type, or the absence of usable result objects, without including search queries, credentials, or payload values.
GroundRoute web search and fetch contain malformed HTTP-success payloads the same way: a non-object payload, a
resultscontainer that is not a list, or a non-empty list with no object entries returns the provider's format error, while non-object entries in an otherwise valid list are skipped in order. Missing, null, or empty results keep the existing "No results found" response.Tavily web search and fetch contain malformed HTTP-success payloads the same way: a non-object payload, or a
results/failed_resultscontainer that is not a list, returns the provider's format error, and non-object entries in an otherwise valid list are skipped in order. Missing, null, or empty results keep the existing response ([]for search,Error: No results foundfor fetch); missing or null search fields (title,url,content) normalize to empty text, and a non-string fetchraw_contentis coerced to text before the truncation limit.Jina, Browserless, and InfoQuest web fetches resolve relative links and image sources using the requested page URL (or a usable HTML base URL), so returned Markdown includes complete destinations. Link resolution preserves the surrounding HTML source, including malformed-page formatting.
Self-hosted Browserless, Crawl4AI, Firecrawl, and fastCRW fetch backends require outbound isolation from private, loopback, link-local, and cloud-metadata networks before setting
network_isolation_confirmed: trueon each tool. DeerFlow rejects private or unverifiable backends by default because those services handle redirects, DNS resolution, and subresources themselves.allow_private_addressespermits intentional internal target URLs and does not bypass the backend check. See delegated fetch deployment;make doctorreports backend configurations that would be refused.Jina fetches support opt-in bounded retries via
max_retries(default0) andretry_budget_seconds(default30) in the tool configuration. ValidRetry-Afterhints set a minimum wait for HTTP 429/503; 429 without a valid hint stays terminal. Hints that cannot fit the remaining budget stop retries. Local backoff remains randomized. Retries may increase upstream requests and cost; see Jina fetch retries.Jina also accepts an opt-in
max_response_bytestool setting (positive integer; omitted/null disables it). It stops oversized decoded responses before extraction, with no partial success or retry. This leaves the 4096-character output cap unchanged and does not bound HTTPX decompressor allocations or wire bandwidth; see response budget.Jina can also bound concurrent HTTP attempts and FIFO waiting with the opt-in
request_admissiontool extra. One immutable budget is shared across clients and agent threads in each process; replicas have independent budgets. See Jina request admission.Run
make doctorat any time to verify your setup and get actionable fix hints. If you are opening a GitHub issue about a local setup or runtime problem, runmake support-bundle. The command prints reporter next steps, writes a*-issue-summary.mdfile to paste into the issue, a*-issue-draft.mdfile for AI-assisted issue filing, and an optional evidence zip under.deer-flow/support-bundles/. If an AI assistant files the issue, start from the draft and replace every REQUIRED placeholder instead of inventing missing facts. Attach the zip only if a maintainer asks for it, or if the summary alone is not enough. Maintainers and AI triage tools can start withtriage.json; the bundle includes redacted diagnostics and file manifests only, and does not include.env, raw conversation messages, or user file contents. Subprocess diagnostics are captured as UTF-8, with Python helpers emitting UTF-8 even on non-UTF-8 hosts and escaping unencodable characters. Thread manifests follow the local launcher's runtime paths: checkout.envvalues override shell exports, and a project-root override alone still usesbackend/.deer-flowfirst. Simple variable references such asDEER_FLOW_HOME="$PWD/backend/.deer-flow"use the checkout asPWD; single-quoted references remain literal. If the checkout.envcannot be read or decoded as UTF-8, thread diagnostics retain shell exports and default storage paths. For standalone Gateway launches usingbackend/.envorDEER_FLOW_ENV_FILE, export the effectiveDEER_FLOW_HOMEwhen collecting the bundle and ensure the checkout.envdoes not override it. Doctor's internal tool probes also decode UTF-8 with replacement for invalid bytes so the remaining diagnostic output stays available.When a thread manifest is requested, a nonempty
DEER_FLOW_HOMEselects the runtime data directory instead of stale checkout data. External data paths appear as{DEER_FLOW_HOME}in the report; file contents remain excluded. Runtime path settings also come from the checkout.env, with shell exports taking precedence. Relative paths resolve from the checkout as inmake dev. Without an explicit home, a configuredDEER_FLOW_PROJECT_ROOTselects its.deer-flowdirectory; external paths use that variable name in the report.Advanced / manual configuration: If you prefer to edit
config.yamldirectly, runmake configinstead to copy the full template. Optional dependency auto-detection accepts UTF-8 configuration files with or without a byte-order mark (BOM). Seeconfig.example.yamlfor the complete reference including CLI-backed providers (Codex CLI, Claude Code OAuth), OpenRouter, Responses API, subagent runtime caps such assubagents.max_total_per_run, and more.Optional per-model pricing must use one currency across all priced models. DeerFlow disables Console cost estimates when currencies are mixed rather than presenting an invalid aggregate.
Administrators can also open Settings → Models to add, edit, test, and enable/disable shared OpenAI-compatible Chat Completions models without editing
config.yaml. Enter a unique name, base URL, model ID, and optional API key; saving refreshes the chat model list. Connection testing sends a short streaming tool-call request and may incur provider charges. It does not save the draft or verify image support; set image support and token limits from provider documentation. Official DeepSeek models athttps://api.deepseek.comorhttps://api.deepseek.com/v1(default HTTPS port) automatically use DeerFlow's DeepSeek adapter, preserving reasoning content across tool calls and honoring output token limits. Chat uses the selected thinking mode; the connection test temporarily disables thinking because DeepSeek rejects forced tool selection in thinking mode. The test checks streaming tool connectivity, not every agent workflow or thinking-mode behavior. Existing saved DeepSeek profiles receive this adapter without re-entering credentials. DeepSeek-specific settings for third-party proxies, other native adapters, and advanced reasoning settings remain YAML-configured.DeepSeek regression tests run offline with the normal backend suite. To verify the real provider explicitly, set
DEEPSEEK_TEST_API_KEYin your environment and run frombackend/:DEER_FLOW_RUN_LIVE_TESTS=1 uv run --no-sync pytest tests/test_managed_deepseek_live.py -qThese opt-in tests send short requests to DeepSeek and may incur charges; they use temporary state, never save credentials to the deployment catalog, and are skipped in CI.
DEEPSEEK_TEST_MODELoptionally selects a different DeepSeek model ID (default:deepseek-flash). The same tests can be run on unfixed and fixed revisions; success is always the expected result.YAML models remain read-only in this page and take precedence on name conflicts. Managed models are appended after YAML models; edits apply to new configuration snapshots, while active runs retain their existing snapshot. Disabling a model removes it from future selection/resolution, so update any custom-agent or scheduled task definitions that explicitly reference it before disabling it. Managed models are shared by the deployment, not personal API-key profiles, and remain subject to the existing model authorization policy.
The encrypted catalog and a generated local encryption key are stored in
$DEER_FLOW_HOME/managed-models/(default.deer-flow/managed-models/). Persist and back up the whole directory, restrict filesystem access, and share it across Gateway workers/replicas that should use the same catalog. The local key is protected by filesystem permissions; encryption does not protect against someone who can read both files. Losing the key requires restoring the backup. Reads and writes fail if the catalog cannot be decrypted, rather than replacing it. This storage is independent of the SQL backend and works with read-only YAML mounts.When several models are configured, open either model picker and use the star beside a model to favorite it. Favorites appear first in both the main chat and Side Chat pickers without changing either chat's selected or default model. They are stored for the signed-in user in the current browser, so they do not sync to another browser or device and do not require a startup setting. The compact favorites picker intentionally omits search and only adds favorite ordering to the two-line model list.
Manual model configuration examples
models:
- name: gpt-4o
display_name: GPT-4o
use: langchain_openai:ChatOpenAI
model: gpt-4o
api_key: $OPENAI_API_KEY
- name: openrouter-gemini-2.5-flash
display_name: Gemini 2.5 Flash (OpenRouter)
use: langchain_openai:ChatOpenAI
model: google/gemini-2.5-flash-preview
api_key: $OPENROUTER_API_KEY
base_url: https://openrouter.ai/api/v1
- name: opper-claude-sonnet-4-6
display_name: Claude Sonnet 4.6 (Opper)
use: langchain_openai:ChatOpenAI
model: claude-sonnet-4-6
api_key: $OPPER_API_KEY
base_url: https://api.opper.ai/v3/compat
- name: gpt-5-responses
display_name: GPT-5 (Responses API)
use: langchain_openai:ChatOpenAI
model: gpt-5
api_key: $OPENAI_API_KEY
use_responses_api: true
output_version: responses/v1
- name: qwen3-32b-vllm
display_name: Qwen3 32B (vLLM)
use: deerflow.models.vllm_provider:VllmChatModel
model: Qwen/Qwen3-32B
api_key: $VLLM_API_KEY
base_url: http://localhost:8000/v1
supports_thinking: true
when_thinking_enabled:
extra_body:
chat_template_kwargs:
enable_thinking: true
OpenRouter and similar OpenAI-compatible gateways should be configured with langchain_openai:ChatOpenAI plus base_url. If you prefer a provider-specific environment variable name, point api_key at that variable explicitly (for example api_key: $OPENROUTER_API_KEY).
The write_file tool's output-budget hint uses the active model's effective max_tokens. Missing or unusable limits, including YAML .inf, -.inf, .nan, and values too large for the character estimate, omit this hint without failing tool assembly; provider-specific model validation still applies.
To route OpenAI models through /v1/responses, keep using langchain_openai:ChatOpenAI and set use_responses_api: true with output_version: responses/v1.
Models whose provider contract differs from DeerFlow's generic thinking/effort assumptions can declare a per-model mapping-valued reasoning: block (thinking unsupported/optional/required, the accepted effort values with aliases and a default, the payload dialect, and the reasoning-history requirement). The setup wizard's Z.AI GLM-5.3-Flash profile uses it: thinking stays on for every foreground and background call, and the effort selector offers the model's own low/high/max levels. Ollama's existing boolean reasoning: true remains a native provider setting and is forwarded to ChatOllama. When migrating a profile to a custom effort path, remove any old reasoning_effort setting from the profile and thinking templates; configuration validation rejects the leftover key. The chat UI drops a remembered provider-specific effort when switching to a legacy model that does not advertise it. Profiles without the block keep their existing provider behavior. See config.example.yaml for the shape and the equivalent manual configuration.
For vLLM 0.19.0, use deerflow.models.vllm_provider:VllmChatModel. For Qwen-style reasoning models, DeerFlow toggles reasoning with extra_body.chat_template_kwargs.enable_thinking and preserves vLLM's non-standard reasoning field across multi-turn tool-call conversations. Legacy thinking configs are normalized automatically for backward compatibility, without modifying model defaults or caller-owned extra_body dictionaries; a reused dictionary can change the legacy switch between requests. If the endpoint reports a cumulative usage snapshot on every streaming chunk, set cumulative_stream_usage: true so DeerFlow converts those snapshots into per-chunk deltas; the option is disabled by default and leaves usage unchanged when a stable completion id is unavailable. Reasoning models may also require the server to be started with --reasoning-parser .... If your local vLLM deployment accepts any non-empty API key, you can still set VLLM_API_KEY to a placeholder value.
When prompt caching is enabled for a model configured with
deerflow.models.claude_provider:ClaudeChatModel, DeerFlow preserves thinking
and redacted-thinking history without placing cache breakpoints directly on
those blocks. Extended-thinking tool follow-ups can keep using prompt caching.
CLI-backed provider examples:
models:
- name: gpt-5.4
display_name: GPT-5.4 (Codex CLI)
use: deerflow.models.openai_codex_provider:CodexChatModel
model: gpt-5.4
supports_thinking: true
supports_reasoning_effort: true
- name: claude-sonnet-4.6
display_name: Claude Sonnet 4.6 (Claude Code OAuth)
use: deerflow.models.claude_provider:ClaudeChatModel
model: claude-sonnet-4-6
max_tokens: 4096
supports_thinking: true
ClaudeChatModel normalizes manual extended-thinking budgets when auto_thinking_budget: true (the default). An omitted or null budget gets 80% of max_tokens, with a minimum of 1024 tokens. Explicit budgets must be integers of at least 1024 and, for ordinary thinking, strictly below max_tokens; output limits of 1024 or less fail locally. For supported manual interleaved thinking, configure tools and betas: ["interleaved-thinking-2025-05-14"]: the thinking budget may equal or exceed the positive output limit. See Anthropic's interleaved-thinking rules for supported models. auto_thinking_budget: false bypasses this normalization and validation; disabled and adaptive thinking are unchanged.
- Codex CLI reads
~/.codex/auth.json - Codex function tools preserve explicit
strict: trueorstrict: falsein wrapped or flat dictionary definitions. Missing or null settings keep the provider default.bind_toolsapplies the same conversion to dictionaries andBaseToolschemas. - Completed Codex responses still return their text and tool calls when token usage is null, omitted, or empty; usage metadata remains unavailable.
- The Codex model provider returns completed responses without waiting for the SSE connection to close. Failed or incomplete responses report the provider's error or reason; partial output is not returned as a successful answer. Non-object error details are reported as text.
- Claude Code accepts
CLAUDE_CODE_OAUTH_TOKEN,ANTHROPIC_AUTH_TOKEN,CLAUDE_CODE_CREDENTIALS_PATH, or~/.claude/.credentials.json CLAUDE_CODE_OAUTH_TOKEN_FILE_DESCRIPTORaccepts a UTF-8 token handoff and reuses it for later model instances in the same process. Undecodable handoffs are skipped so Claude Code can still try its override or default credentials file.- CLI credential JSON files accept UTF-8 with or without a BOM, independently of the host locale.
make doctoraccepts the same files when checking CLI authentication. Invalid text encoding is treated as an unreadable source; Claude Code can still try its default file after an invalid override. - ACP agent entries are separate from model providers — if you configure
acp_agents.codex, point it at a Codex ACP adapter such asnpx -y @zed-industries/codex-acp - A bare or absolute ACP agent
commandis resolved before the agent starts, so npm-installed launchers work on Windows too:npxormcodeis spawned as itsnpx.cmd/mcode.cmdshim instead of failing with a command-not-found error, and an absolute path such asC:\Users\me\AppData\Roaming\npm\mcoderesolves to themcode.cmdshim beside npm's extensionless shell script instead of failing withWinError 193. The lookup uses thePATHthe agent subprocess will actually see (acp_agents..env.PATHwhen set, otherwise the Gateway's) and hands the spawn an absolute path, so a relativePATHentry cannot be re-interpreted inside the agent's own workspace. A relativecommand(for examplebin/mcode) is used as configured, and a command that cannot be resolved is still reported with the install guidance below. - Each ACP agent's
timeout_seconds(default: 1800) is one shared budget for initialization, session creation, and the prompt, starting after the subprocess launches. On timeout, DeerFlow aborts the invocation and closes the subprocess before returning an error. Workspace/MCP preparation and subprocess cleanup are outside this budget. ATimeoutErrorraised by the SDK before this deadline expires retains its own error message. - MiniMax Code speaks ACP directly. Install and authenticate it, then add it as an ACP agent:
npm install --global @minimax-ai/code
mcode login
acp_agents:
mcode:
command: mcode
args: ["acp"]
description: MiniMax Code for implementation, refactoring, debugging, and repository tasks
auto_approve_permissions: false
mcode must be on the Gateway process's PATH; installing it only on the Docker host does not make it available inside the Gateway container. DeerFlow invokes it through invoke_acp_agent in a per-thread ACP workspace and forwards enabled MCP servers. Keep auto_approve_permissions: false for untrusted tasks; enable it only when mcode must edit files or run commands and you trust the task.
- On macOS, export Claude Code auth explicitly if needed:
eval "$(python3 scripts/export_claude_code_oauth.py --print-export)"
The exporter rejects malformed credential containers and non-string or blank access tokens before printing a token, emitting a shell export, or writing a credentials file. Valid tokens are exported unchanged, and file exports preserve the full credential container.
API keys can also be set manually in .env (recommended) or exported in your shell:
OPENAI_API_KEY=your-openai-api-key
TAVILY_API_KEY=your-tavily-api-key
For an explicit backend dotenv file, export DEER_FLOW_ENV_FILE before startup,
alongside DEER_FLOW_CONFIG_PATH if needed. For example, from backend/:
DEER_FLOW_ENV_FILE=/srv/deer-flow/stage.env DEER_FLOW_CONFIG_PATH=/srv/deer-flow/stage.yaml make gateway
Relative dotenv paths use the backend process working directory; absolute paths
work regardless of that directory. Existing process variables win. An unset
selector preserves default dotenv discovery; a specified empty, missing,
non-file or unreadable path fails startup. Explicit selection also fails when
PYTHON_DOTENV_DISABLED disables dotenv loading. Restart after changing the file.
This selects backend dotenv input only: shell launchers, Docker Compose and the
frontend retain their own environment loading. Values they already export win.
It does not select ENV profiles or isolate databases, storage or tenants.
See backend dotenv selection.
Running the Application
Deployment Sizing
Use the table below as a practical starting point when choosing how to run DeerFlow:
| Deployment target | Starting point | Recommended | Notes |
|---|---|---|---|
Local evaluation / make dev |
4 vCPU, 8 GB RAM, 20 GB free SSD | 8 vCPU, 16 GB RAM | Good for one developer or one light session with hosted model APIs. 2 vCPU / 4 GB is usually not enough. |
Docker development / make docker-start |
4 vCPU, 8 GB RAM, 25 GB free SSD | 8 vCPU, 16 GB RAM | Image builds, bind mounts, and sandbox containers need more headroom than pure local dev. |
Long-running server / make up |
8 vCPU, 16 GB RAM, 40 GB free SSD | 16 vCPU, 32 GB RAM | Preferred for shared use, multi-agent runs, report generation, or heavier sandbox workloads. |
- These numbers cover DeerFlow itself. If you also host a local LLM, size that service separately.
- Linux plus Docker is the recommended deployment target for a persistent server. macOS and Windows are best treated as development or evaluation environments.
- If CPU or memory usage stays pinned, reduce concurrent runs first, then move to the next sizing tier.
Option 1: Docker (Recommended)
Requires Docker Desktop / Docker Engine and Docker Compose v2.24+
(docker compose version). Older Compose clients cannot parse the optional
env_file syntax in docker/docker-compose.yaml and docker/docker-compose-dev.yaml.
Development (hot-reload, source mounts):
make docker-init # Pull sandbox image (only once or when image updates)
make docker-start # Start services (auto-detects sandbox mode from config.yaml)
make docker-logs # View logs
make docker-start starts provisioner only when config.yaml uses provisioner mode (sandbox.use: deerflow.community.aio_sandbox:AioSandboxProvider with provisioner_url).
Docker builds use the upstream uv registry by default. If you need faster mirrors in restricted networks, export UV_INDEX_URL=https://pypi.tuna.tsinghua.edu.cn/simple and NPM_REGISTRY=https://registry.npmmirror.com before running make docker-init or make docker-start.
Local AIO sandbox control traffic is always direct: loopback/private addresses,
single-label cluster hosts, and Docker/Podman internal hostnames do not inherit
HTTP_PROXY or HTTPS_PROXY. External sandbox FQDNs and public IPs still
honor environment proxy settings.
Backend processes automatically pick up config.yaml changes on the next config access, so model metadata updates do not require a manual restart during development.
Gateway runs use the top-level recursion_limit in config.yaml when an API
request does not provide one. The default is 100; valid per-request values
take precedence, and max_recursion_limit (default 1000) caps both. Changes
apply to the next run without restarting the Gateway. This top-level setting
applies to Gateway API runs; IM channel and embedded DeerFlowClient runs
retain their own defaults and per-call override paths.
The checkpoint storage settings database.checkpoint_channel_mode and
database.checkpoint_delta.snapshot_frequency (default 10) are exceptions:
both are frozen when the process first builds an agent (including through
DeerFlowClient) and require a process restart to change safely.
The optional database.checkpoint_cache section (delta channel mode only)
caches materialized checkpoint histories: type is memory (default) or
redis, and max_entries: 0 disables the cache. The redis backend is
Gateway/async-only; the sync TUI/embedded path supports memory only. The
cache is performance-only — results are identical with it disabled — so it is
never frozen and workers sharing one checkpoint database may safely run
different cache settings.
[!TIP] On Linux, if Docker-based commands fail with
permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock, add your user to thedockergroup and re-login before retrying. See CONTRIBUTING.md for the full fix.
Production (builds images locally, mounts runtime config and data):
make up # Build images and start all production services
make down # Stop and remove containers
Access: http://localhost:2026
make up waits for the Gateway /health endpoint before reporting success.
If the Gateway does not become healthy within the startup window, deployment
exits non-zero and prints the container status plus recent Gateway logs. The
production image starts from its already-built environment and never resolves
or installs Python dependencies at container startup.
If make up reports an unwritable runtime home or an unreadable persisted secret
after make docker-start, run the printed sudo chown -R : ''
recovery command and retry. The check honors secret overrides from the shell or
.env, accepts readable read-only secret files, and leaves make down available
without reading or generating secrets.
For persistent deployments, configure database.backend as sqlite or
postgres. The selected backend is shared by the LangGraph checkpointer,
LangGraph Store, and DeerFlow application data. The deprecated checkpointer
section, when present, overrides the first two for backward compatibility.
Gateway startup automatically repairs the missing run-change schema affecting some existing databases (#5516). The repair preserves run history and existing change positions; downgrading the repair to its predecessor also retains the schema and positions required by that version.
For lightweight single-process event persistence, run_events.backend: jsonl
keeps Unicode message content intact, including line and paragraph separators.
Existing valid JSONL records remain readable without rewriting the files.
For per-call usage audits, an immediate LLM response replay with populated usage updates both nested usage and top-level token counters, even if initial usage was absent or zero. The original request, model, and status metadata stay intact; flushed events are not rewritten. See the run event contract.
The unified nginx endpoint is same-origin by default and does not emit browser CORS headers. If you run a split-origin or port-forwarded browser client, set GATEWAY_CORS_ORIGINS to comma-separated exact origins such as http://localhost:3000; the Gateway then applies the CORS allowlist and matching CSRF origin checks.
When fine-grained authorization is enabled, Live Browser connections require threads:write as well as ownership of the thread, even when only viewing frames: the same connection can control the browser. Permission checks run when connecting. Restart Gateway after upgrading to disconnect sessions admitted by older code.
Browser login uses HttpOnly session cookies. The login page offers a "keep me signed in" option that extends the browser session when the request is HTTPS (including trusted X-Forwarded-Proto: https) or localhost HTTP. The localhost exception uses the direct request Host and ignores forwarded host headers. Public HTTP deployments, including many temporary sandbox URLs, fall back to session cookies by default. DeerFlow never stores the password in browser storage; the UI may remember only the email address.
Administrators can suspend and restore accounts from Settings → Users (admin-only; the section is hidden for ordinary users). The page lists each account's email, role, and status and exposes the same enable/disable operation as PATCH /api/v1/admin/users/{id}; the last remaining active admin cannot be disabled, and the Gateway's rejection detail is shown verbatim. A disabled account is rejected by every authentication surface, and its live sessions end on the next request with a login page that states the account has been disabled. Password login for a disabled account deliberately still answers "incorrect email or password" so the state is not disclosed to whoever holds the credentials.
DeerFlow still uses Forwarded / X-Forwarded-* headers to recover the browser-facing scheme and origin behind a proxy. The bundled nginx sets X-Forwarded-Proto, but preserves an upstream HTTPS value and does not overwrite every forwarded header. Configure the outer trusted proxy to replace or strip client-supplied forwarding headers before traffic reaches DeerFlow.
[!IMPORTANT] The Gateway still owns active run tasks in process, so production defaults to a single Gateway worker (
GATEWAY_WORKERS=1). Multi-worker deployments require Postgres, the Redis stream bridge (stream_bridge.type: redis),run_ownership.heartbeat_enabled: true, andrun_events.backend: db; process-local memory/JSONL event stores cannot enforce singleton delivery receipts across workers. Kubernetes replicas run one worker per Pod, which the worker count cannot see: declare them withdeployment.multi_instance: true(orDEER_FLOW_MULTI_INSTANCE=1, which deploy tooling such as a Helm chart can set from its replica count) so the startup gate enforces the same prerequisites instead of staying inert. Declared instances that store credentials (channel_connections) must also share one at-restDEER_FLOW_CREDENTIALS_KEY(the Helm chart andmake upgenerate it; back it up, see backend/docs/CONFIGURATION.md). The bridge shares SSE delivery and boundedLast-Event-IDreplay across workers. When a valid reconnect cursor has been trimmed, or a subscriber that already established an empty-stream wait falls behind before its first delivery, Memory and Redis emit a machine-readable SSEgapevent instead of silently returning a partial replay; the Web UI reloads durable thread/event state and resumes from the retained tail. Lease reconciliation marks runs from dead workers as errors, persists their delivery receipts, publishes the terminal stream marker, schedules retained-stream cleanup, and updates the affected thread status. SSE,/wait, and internal stream consumers usestream_bridge.heartbeat_interval_seconds(default15) for idle liveness checks; changing it requires a Gateway restart. Malformed Redis reconnect IDs live-tail new events instead of replaying the retained buffer, and the rolling retained-buffer TTL (stream_ttl_seconds) remains a cleanup safety net rather than a run timeout. Failed-login counters and lockouts forPOST /api/v1/auth/login/local(auth.local.max_login_attempts/lockout_seconds) are kept in the sharedlogin_throttletable whenever the application database is SQLite or Postgres (auth.local.throttle_storage: auto, the default), so every replica enforces one lockout per client IP;memorykeeps the historical per-process counter, which under N replicas hands an attacker N ×max_login_attemptsguesses and logs a startup warning. IM chat-to-thread bindings live in the sharedchannel_thread_bindingstable whenever the application database is SQLite or Postgres (an existingchannels/store.jsonis imported once at startup and renamedstore.json.migrated); the remaining IM channel state (each instance's platform connections, follow-up buffers) still needs its own multi-worker coordination. To try this topology on one machine,make dev-multistarts two Gateways on shared Postgres, Redis andDEER_FLOW_HOME, andmake dev-multi-checkruns the cross-instance checks; see the local two-Gateway harness in backend/docs/CONFIGURATION.md.During SSE replay-gap recovery, the Web UI preserves a reconnect pointer claimed by another run while durable state is loading. An older recovery cannot overwrite that pointer when it rejoins its own stream. Even a stale foreign pointer is preserved, so the next reload may first check that run before active-run recovery rejoins, adding a reconnect cycle.
Store-only SSE and
/waitconsumers can finish after a heartbeat when run-manager or scheduled-task reconciliation recovers a dead worker's run, even if no END marker was published. Ordinary completions still wait for END to preserve tail events.In single-process JSONL deployments, cancelling an admitted event-store mutation waits for its background file I/O, rollback, and bookkeeping to settle before releasing the thread write lock. This prevents an older cancelled write from recreating deleted records or rolling back a later successful write. Cancellation can therefore wait on slow storage; it does not stop an in-flight filesystem operation. Callers still waiting to acquire the lock can cancel without starting a mutation. A batch spanning multiple threads drains its current thread group before propagating cancellation; subsequent thread groups do not start.
After a run publishes its terminal stream marker, its process-local
RunRecordremains available for the existing five-minute grace period before cleanup; durable run history remains available throughRunStore, while the stream bridge retains its delivery tail on its separate cleanup schedule.Run cancellation may land on any Gateway worker. A non-owning worker now persists the interrupt or rollback request for the live owner, which observes it during lease renewal and performs the normal cancellation flow; load-balancer routing alone no longer produces a 409. The first accepted action wins even if a retry lands on the owner, and accepted cancellation competes atomically with owner completion. Dead owners still follow lease takeover and orphan recovery. Cancellation latency is therefore bounded by the lease heartbeat interval.
Cancelling a model recovery probe, including while it is queued or waiting to retry, lets the next call check whether the provider has recovered. Cancellation does not count as a provider failure or release another call's active recovery probe.
Model retries honor valid provider
Retry-Afterhints up to 24 hours, including waits longer than the local backoff cap. Malformed, non-finite, or over-limit hints use normal backoff instead. This applies to numeric/date hints in LLM error handling and integer-second hints in Claude's retry path.With lease heartbeat enabled, a transient RunStore renewal error is retried only until the last confirmed lease expires; the stale worker then cancels local execution and suppresses checkpoint, completion-hook, delivery-receipt, and thread-status finalization. A remote tool side effect already in flight may still be outside local cancellation.
Reconciliation uses an atomic takeover claim that re-checks the lease after candidate selection, so a successful owner renewal wins over orphan recovery and only one reconciler can report a run as recovered. When multiple Gateway workers share the Docker/AIO or E2B sandbox backend, also configure
sandbox.ownership.type: redis; E2B uses the leases during background startup and periodic reconciliation so duplicate/orphan cleanup cannot terminate a live peer's sandbox.
See CONTRIBUTING.md for detailed Docker development guide.
Upgrading an existing checkout
Keep config.yaml, .env, and extensions_config.json. Stop the services you
currently use, run git pull --ff-only, then start the same mode again. Do not run
make config or make docker-init again for a routine source upgrade. If the new
version requires configuration changes, run make config-upgrade before restarting.
See Operations and Troubleshooting
for the commands for each mode.
When rolling out the resume-command idempotency fix across multiple Gateway
workers or Pods, route all keyed resume submissions and retries (command.resume
with Idempotency-Key, on thread-scoped /runs, /runs/stream, or /runs/wait)
only to upgraded workers until every worker serving these endpoints is upgraded.
Older workers ignore the private resume identity and compare only input: with
input: null, retrying deny can incorrectly reuse an earlier approve run.
If the load balancer cannot isolate upgraded workers, pause keyed resume traffic
until the rollout finishes, or stop all old workers before restarting on the new
version. The additive database migration keeps old run-history readers compatible;
it does not make old workers safe for resume admission. An upgraded worker returns
409 when retrying an identity-less legacy resume run, even for the same decision;
inspect that run and the current thread state before deciding to submit a new
action. See the run API contract.
Custom run stores must explicitly support every supplied idempotency field.
Keyed input and resume requests return HTTP 503 before run admission when the
store cannot guarantee that support. The built-in Memory and SQL stores support
both fields; older custom stores with an explicit idempotency_key parameter
continue to accept keyed input requests.
Option 2: Local Development
If you prefer running services locally:
Prerequisite: complete the "Configuration" steps above first (make setup). make dev requires a valid config.yaml in the project root. Set DEER_FLOW_PROJECT_ROOT to define that root explicitly, or DEER_FLOW_CONFIG_PATH to point at a specific config file. Runtime state defaults to .deer-flow under the project root and can be moved with DEER_FLOW_HOME; skills default to skills/ under the project root and can be moved with DEER_FLOW_SKILLS_PATH. Run make doctor to verify your setup before starting.
Configuration checks and agent-storage migrations use the installed backend environment through uv run --no-sync --project backend from the repository root. This preserves relative runtime paths and installed optional dependencies; see the setup checks and agent-storage migration.
On Windows, run the local development flow from Git Bash. Native cmd.exe and PowerShell shells are not supported for the bash-based service scripts, and WSL is not guaranteed because some scripts rely on Git for Windows utilities such as cygpath.
The documented root make commands invoke repository .sh files through Bash
explicitly. They therefore continue to work from source archives or filesystems
that do not preserve POSIX executable bits. When calling a script directly from
such a checkout, use bash ./scripts/.sh ....
-
Check prerequisites:
make check # Verifies Node.js 22+, pnpm, uv, nginxThe local
make check,make install,make dev, andmake startentry points use a directpnpmexecutable when available and otherwise fall back tocorepack pnpm. With native Windows Python, the shared runner checkspnpm.cmdbefore the genericpnpmlookup, which followsPATH/PATHEXTand may select an.exeor.batin the same or an earlier PATH directory. The Corepack fallback likewise checkscorepack.cmdbeforecorepack. POSIX Python keeps the generic names first, including when running under MSYS/Cygwin. The runner and diagnostics resolve repository paths absolutely, so these checks work regardless of the caller's current directory. Corepack runs fromfrontend/, so it honors thepackageManagerversion pinned infrontend/package.json; enabling a global pnpm shim is not required.make checkdecodes tool diagnostics as UTF-8 and preserves Unicode pnpm errors on non-UTF-8 hosts. Invalid output bytes are replaced so the checker can report dependency failures instead of aborting while decoding output. -
Install dependencies:
make install # Install backend + frontend dependencies + pre-commit hooksHook setup calls pre-commit through uv, so uv's tool directory need not be on
PATH. -
(Optional) Pre-pull sandbox image:
# Recommended if using Docker/Container-based sandbox make setup-sandboxReads the configured sandbox image from UTF-8
config.yaml, with or without a leading BOM, using LF or CRLF line endings. On macOS, a successful Apple Container pull completes this step even when Docker is not installed. If Docker is available, its image is also pulled. -
Start services:
make dev -
Access: http://localhost:2026
-
(Optional) Load sample memory data for local review: open
Settings > Memory, click Import memory, and selectbackend/docs/memory-settings-sample.json. The browser imports into the signed-in user's memory.To replace memory for every registered user in a disposable review environment:
cd backend uv run python ../scripts/load_memory_sample.py --all-usersBulk mode supports SQLite/PostgreSQL user registries, creates timestamped backups under
.deer-flow/memory-sample-backups/, and rejects the non-persistentdatabase.backend: memorymode. See backend/docs/MEMORY_SETTINGS_REVIEW.md for the complete review flow.
Local services always use their internal ports (8001, 3000, and 2026).
The root .env variable PORT configures only the published Docker ingress;
it does not change the Next.js port used by make dev.
Startup Modes
DeerFlow runs the agent runtime inside the Gateway API. Development mode enables hot-reload; production mode uses a pre-built frontend.
| Local Foreground | Local Daemon | Docker Dev | Docker Prod | |
|---|---|---|---|---|
| Dev | ./scripts/serve.sh --dev |
|||
make dev |
./scripts/serve.sh --dev --daemon |
|||
make dev-daemon |
./scripts/docker.sh start |
|||
make docker-start |
— | |||
| Prod | ./scripts/serve.sh --prod |
|||
make start |
./scripts/serve.sh --prod --daemon |
|||
make start-daemon |
— | ./scripts/deploy.sh |
||
make up |
| Action | Local | Docker Dev | Docker Prod |
|---|---|---|---|
| Stop | ./scripts/serve.sh --stop |
||
make stop |
./scripts/docker.sh stop |
||
make docker-stop |
./scripts/deploy.sh down |
||
make down |
|||
| Restart | ./scripts/serve.sh --restart [flags] |
./scripts/docker.sh restart |
— |
make start and make start-daemon rebuild the frontend with next build on
every run. To reuse the last build instead, pass SKIP_FRONTEND_BUILD=1 (or add
--skip-frontend-build when calling ./scripts/serve.sh --prod directly). This
is opt-in: it fails fast when frontend/.next has no completed build.
Gateway owns /api/langgraph/* and translates those public LangGraph-compatible paths to its native /api/* routers behind nginx.
Cold agent imports during run creation use a dedicated worker pool, keeping unrelated Gateway requests responsive while the agent stack loads. A failed factory import prevents the run from being admitted.
For a read-only demo without the Gateway, run make build-static from frontend/,
then HOSTNAME=127.0.0.1 PORT=3000 node --env-file=.env .next/standalone/server.js
from the same directory. The build includes public demo assets and resolves
supported demo API reads locally; writes are unavailable. To display the homepage
GitHub star count, set GITHUB_OAUTH_TOKEN in frontend/.env before starting Node.
The token stays on the server; missing credentials or GitHub failures hide the
count. Restart Node after changing the token; no rebuild is needed.
LangGraph Studio (Optional)
The default make dev topology uses DeerFlow's Gateway-embedded runtime and
does not require LangGraph Studio. To inspect and test the registered lead-agent
graph with the standalone development server, run the command from backend/
so the CLI discovers langgraph.json:
cd backend
uv run langgraph dev --allow-blocking
The command prints the local API and Studio UI URLs. This in-memory server is
for development and testing only. The flag permits DeerFlow's synchronous
configuration and graph-factory setup during local Studio requests; it must not
be treated as a production-server setting. Local Studio authentication is
handled automatically, so the connection does not require custom headers. Use
DeerFlow's documented production startup modes or a supported LangSmith
deployment for production workloads. Assistant ownership and provenance in this
standalone mode are server-owned: Studio can discover registered graphs and the
assistants it creates, and normal assistant-version selection remains available.
Before the locked local runtime loads its persisted development store, DeerFlow
repairs legacy assistant rows and version history so historical client metadata
cannot restore server privileges or be discarded by the runtime's startup
cleanup. Keep the backend dependencies synchronized with uv sync; this
compatibility path requires the declared LangGraph runtime versions and logs a
warning if the persisted-store contract no longer matches its expectations.
The documented command uses LangGraph's file-based custom-app loader, which is
also covered directly by DeerFlow's regression tests.
Standalone runs using if_not_exists="create" retain config and run metadata
on the newly created thread, including searchable tags; run metadata takes
precedence for duplicate keys. Thread ownership and MCP incarnation remain
server-owned, and later runs do not replace the thread's creation metadata.
For workflows that invoke backend/langgraph.json through LangGraph Studio or
a direct LangGraph Server, DeerFlow consumes the authenticated identity
published by that runtime and uses it for custom-agent configuration/SOUL, user
skills and skill policy, uploads, thread data, and memory reads/writes. This
keeps authenticated runs out of the shared default filesystem bucket, and the
server-owned identity takes precedence over ordinary client-supplied user_id
values. External identities such as email addresses are mapped to stable,
collision-resistant directory-safe user IDs before accessing DeerFlow storage.
The default DeerFlow service topology remains the Gateway-embedded runtime
described above.
Gateway runs automatically enforce native delivery for artifacts created or modified under /mnt/user-data/outputs: present_files must present at least one output produced by the current run, and the terminal run.delivery receipt must be durably recorded. Virtual artifact paths are resolved within the same authenticated user and thread scope that produced the output before the output-directory boundary is validated. Runs that do not produce output artifacts keep ordinary conversational behavior.
Thread-scoped runs.wait() calls that finish with status: error report the current run error instead of an earlier answer. The asynchronous Python LangGraph SDK raises by default; pass raise_error=False to inspect the returned status and error.
DeerFlow's built-in custom events are available through both LangGraph streaming interfaces: native clients can continue subscribing to stream_mode="custom", while callback-based integrations can consume the same payloads as on_custom_event records from astream_events(version="v2"). The callback event name matches the payload's type field.
Docker Production Deployment
./scripts/deploy.sh supports building and starting separately:
# One-step (build + start)
./scripts/deploy.sh
# Two-step (build once, start later)
./scripts/deploy.sh build # build all images
./scripts/deploy.sh start # start pre-built images
# Stop
./scripts/deploy.sh down
Advanced
Sandbox Mode
DeerFlow supports multiple sandbox execution modes:
- Local Execution (runs sandbox code directly on the host machine)
- Docker Execution (runs sandbox code in isolated Docker containers)
- Docker Execution with Kubernetes (runs sandbox code in Kubernetes pods via provisioner service)
Sandbox references in conversation state are server-owned. External run and
thread-state APIs reject caller-supplied sandbox values; when restoring a
checkpoint, the runtime resolves the reference against the authenticated user
and thread before a tool can reuse it. A missing runtime thread ID raises an
error even when the referenced sandbox is cached.
When host Bash is enabled for Local Execution, DeerFlow starts OS detection with uname -s, then uses sw_vers on Darwin. On Linux, it reads host system files such as /etc/os-release only when the active sandbox policy permits it. Host filesystem path checks still apply; after a blocked path, the agent is directed to use a permitted command-only probe or virtual path instead of repeating the rejected command.
For Docker development, service startup follows config.yaml sandbox mode. In Local/Docker modes, provisioner is not started.
Local AIO sandbox port allocation stops at TCP port 65535. If the remaining ports from the configured starting port are occupied, allocation reports that no port is available instead of attempting an out-of-range bind.
See the Sandbox Configuration Guide to configure your preferred mode.
Remote directory listings report traversal failures (for example, unreadable directories) as incomplete results, even when no entries were returned. A missing start path is reported separately as “Directory not found.”
The optional Tenki cloud sandbox provider
uses Tenki SDK 1.4.0 or newer. Timed-out commands preserve partial output and
report Exit Code: 124; unsuccessful health checks cannot reclaim a warm sandbox.
Health probes tolerate login-shell output around the ok line, and failures log
the sandbox ID and probe output before replacing the sandbox.
BoxLite shutdown rejects late VM registration and keeps its SDK loop open while in-flight acquisitions finish. If they cannot drain within five seconds, shutdown fails with resources still owned and can be retried.
MCP Server
In the chat UI, enable Token Usage → Debug to inspect generic/MCP tool calls. Each Tool details panel starts collapsed and shows the tool name, call ID, input, and received result or explicit error. Large previews are truncated; fields whose names exceed the remaining preview budget are omitted rather than renamed. Structured previews retain complete JSON syntax, including escaped strings and closing delimiters. Array previews stop when the text budget cannot display another element; literal ellipsis values are preserved. Consecutive generated markers at an array's end share one ellipsis indicating an omitted suffix; markers before later values retain their positions. Text results retain their original representation, including large numeric IDs and duplicate JSON keys, without reparsing. Text exceeding the limit is shown as a prefix with a truncation notice; structured objects and arrays are formatted separately. Copy actions copy only the displayed preview. This is a frontend view of data already received by the browser, without an additional secret-redaction layer.
In plan mode, malformed TODO statuses return normal tool-validation errors without aborting token attribution, so the agent can correct the call.
Tool-produced paths and URLs can be retained as short artifact handles across context compaction (tool_artifacts in config.yaml). Handles distinguish separate tool-result occurrences, even when a provider reuses call IDs. Detected file URLs preserve their query strings and fragments. When PII redaction is enabled, model-visible artifact labels follow that policy; internal references stay intact for tool argument resolution. The configured registry limit retains the newest artifacts, while checkpointed processing identities prevent evicted results from being recaptured after restart. Resolution runs before authorization and write-safety checks; unknown or expired handles return an error without executing the tool. Small unknown structured results may be retained as complete JSON up to 4096 UTF-8 bytes; empty or oversized payloads are skipped. Handles are agent-local: task arguments resolve parent handles to concrete references, and delegated reports must return concrete references rather than child-local handles. A truncated model projection reports how many handles are omitted.
DeerFlow supports configurable MCP servers and skills to extend its capabilities.
When durable MCP background tasks are enabled, agents can use list_background_tasks(status="failed") or status="input_required" to find failed tasks or tasks awaiting input in the current chat. Other supported statuses are submitted, working, completed, and cancelled; omitting the status preserves existing behavior. The database applies the filter before limiting results to the 20 most recent matching tasks. When combined with active_only=true, both filters apply: active statuses are submitted, working, and input_required, so terminal statuses return an empty list. GET /api/threads/{thread_id}/mcp-tasks supports the same status and active_only filters before its limit (default 50, range 1–100); for example, ?status=failed&limit=20. Unknown statuses return 422, and omitting the filters preserves the existing response.
For HTTP/SSE MCP servers, OAuth token flows are supported (client_credentials, refresh_token).
Missing, malformed, or out-of-range token response expires_in values use a one-hour default lifetime. This includes lifetimes that cannot be added to the current time without overflowing the expiry timestamp.
Durable HTTP/SSE task status and cancellation calls select configured user_auth credentials using the persisted task owner, including after restart; per-request secrets are not retained for background calls. If a request-scoped credential overrides submit authentication, both credentials must authorize access to the same remote task.
For stdio MCP servers, per-tool call timeouts can be configured with tool_call_timeout; durable background-task calls honor the same setting for HTTP/SSE servers as well.
For stdio file outputs, a bare filename is linked to a uniquely matching file created or changed by that call. Filenames embedded in unrelated paths, including Windows backslash paths, are left intact.
For HTTP/SSE background-task calls, session_init_timeout separately bounds connection setup (including the SSE endpoint event) and MCP initialization together; it stops applying once the tool call begins. Initialization deadline errors identify the server and configured time limit.
Ordinary task subagents retain the parent run's captured thread incarnation for MCP calls, including legacy threads, so delegation preserves the same lifecycle scope.
MCP tool names are prefixed with _ by default to prevent collisions across servers. If a server already namespaces its own tools, set tool_name_prefix: false on that server in extensions_config.json to keep the original names. Disable the prefix only when the resulting names remain unique across all enabled servers.
Signed-in users' notification toggle, default model, conversation mode, and reasoning effort are saved to their account and restored on other browsers or after clearing browser storage. Browser notification permission still needs to be granted on each device. Changes retry after network failures; unsent changes survive a reload in the same tab. Concurrent edits to different fields are preserved; for the same field, the last server write wins. Existing unscoped browser preferences are not uploaded automatically because they have no account owner; reselect those settings once after upgrading. Static demos and auth-disabled development keep browser-local settings. Thread-specific model overrides and other display preferences remain local.
In a new chat, the submitted question stays above its streamed reasoning and tool steps while the server creates the conversation and confirms the message.
Markdown and JSON conversation exports from the chat header or sidebar read all persisted history pages, including earlier turns outside the loaded view or compacted model context. A failed history read stops the export instead of downloading a partial transcript. Public demos export their bundled messages.
Capability Center groups plugins by office collaboration, documents and knowledge, search and research, business and data, and development and operations. The directory includes setup references alongside existing MCP configurations and Lark. Recommended integrations and built-in support do not imply an installed or verified connection; the Installed filter shows configured MCP entries and installed Lark only.
Personal MCP connections configured in the web interface are persisted per user. Deployment tools remain shared. Administrators can add, edit, enable, disable and delete shared MCP servers under Platform provided; ordinary users see their status without controls. Personal plugin switches affect only the signed-in user's connections. See connection ownership.
For plugin manifests, adapter registration, and Agent capability selection, see Capability Center integration contract.
DingTalk and WeCom group notifications and HubSpot CRM are bundled configurable plugins. Administrators supply robot credentials or a HubSpot private app token; Agents can then send requested group notifications, list companies, or create contacts. Saving configuration performs no external write. These plugins reuse the existing MCP lifecycle and require no separate plugin service. See the integration contract above for required fields, scopes, and feature boundaries.
Plugin brand icons are bundled locally. When adding or editing one personal MCP plugin, users can upload a PNG, JPG, or WebP image (up to 2 MB), preview it, or restore the default icon. Changes take effect only after Save; custom icons persist across browsers as a normalized 128px PNG in the server entry's display-only presentation.icon metadata. They are not sent to the MCP transport.
Capability Center > Plugins adds, replaces, and deletes one MCP server at a time through targeted mutations that preserve concurrent sibling changes; deletes use a bodyless URL-addressed request. An invalid stdio command on one server no longer blocks toggling another, while enabling that invalid server remains protected by the command allowlist and surfaces the backend validation message in the UI.
Targeted updates accept both DeerFlow's type field and the MCP-spec transport field for SSE/HTTP servers.
Runtime MCP and skill updates replace extensions_config.json atomically, so an interrupted write cannot leave the shared configuration truncated or partially written. Every Gateway worker or instance that reads the same file picks up a new revision on its next read (the cache checks the file's content signature), so MCP and skill changes made through one replica apply to the others without a restart; a missing, truncated, or invalid revision keeps the previous configuration until a complete one lands, including when the file disappears during a reload.
The admin MCP cache reset advances a durable generation marker in the writable config directory. Every Gateway worker mounting that same directory retires its own cached tools and pooled sessions before the next lookup; replicas with independent filesystems are not implicitly covered. If no config path is available, the API reports a process-local reset instead.
The parsed extensions configuration and its recorded content digest come from the same read, so a racing edit followed by a timestamp-preserving backup restore cannot leave a different revision cached indefinitely.
extensions_config.json accepts UTF-8 with or without a leading byte-order mark (BOM), including files saved as UTF-8 with BOM by an editor.
When deferred MCP tool discovery is enabled, tool_search supports opt-in literal keyword queries such as keywords:notebook jupyter. Keywords match names and descriptions in any order, ranked by distinct term coverage and then name hits; ties retain catalog order. This mode uses the first 256 characters after the prefix and at most 16 unique whitespace-separated terms, returning up to five tools. Punctuation stays literal. Bare queries and +required ranking keep their existing regular-expression semantics; select:tool_a,tool_b still fetches all exact, case-sensitive name matches.
MCP routing hints can also prefer a specific MCP tool for matching requests without forbidding other tools. When tool_search defers MCP schemas, matching routing metadata can auto-promote up to tool_search.auto_promote_top_k deferred schemas before the model call.
Deferred tool_search regex queries share a 100 ms execution budget across the catalog, including +required ranking. Patterns are limited to 256 characters and searchable names/descriptions to 65,536 characters each. Queries that exceed a limit return a retry hint without promoting partial results; use a simpler pattern, keywords:, or select: with exact names. Oversized catalog fields are reported with the tool name, field and length, and logged with the MCP server name when available; use keywords: or select: because simplifying the pattern cannot shrink a catalog field. Exact selection remains uncapped and bypasses regex limits; the literal keywords: mode also bypasses regex limits. Invalid regex syntax still falls back to a literal substring match.
OpenViking users can register the official Streamable HTTP endpoint at /mcp
with an owner-bound USER API key. The native forget tool is exposed for
capability parity; deletion is irreversible, so it should be called only after
explicit user confirmation. DeerFlow does not enforce that confirmation. This
explicit, model-selected MCP tool path can run alongside the separate automatic
OpenViking memory backend; it does not replace automatic turn capture or recall. See the
OpenViking MCP tools configuration.
The Gateway can adapt an MCP server's ordinary submit / status / cancel tools into durable background tasks. The Agent sees only the configured submit tool and a DeerFlow-local task ID; remote IDs are persisted before the submit call returns, while status and cancel stay internal to the runtime. Polling uses cross-worker leases, exponential retry backoff, scoped MCP sessions, bounded result storage, and restart recovery. A status-tool isError is retained as a bounded diagnostic and retried; servers report a permanent remote-task outcome through a normal structured result with status: "failed". Remote poll hints are finite positive numbers capped at 24 hours, artifact-reference JSON is limited to 64 KiB, and task/server identifiers are validated against their durable SQL column limits before persistence. Input-required and terminal updates wake the current chat through idempotent Agent runs, while list_background_tasks and cancel_background_task let the Agent manage tasks without asking users for remote handles. Current-thread tasks are available through GET /api/threads/{thread_id}/mcp-tasks, its detail endpoint, and POST /api/threads/{thread_id}/mcp-tasks/{task_id}/cancel; when the task runtime actually starts, the Web UI exposes the same safe local view from the chat header with live status refresh, cancellation, and on-demand result, artifact, input-request, status-error, and cancellation-retry details. Default-disabled and memory-backend deployments hide that UI and do not poll the task endpoints. A failed remote cancellation remains queued with backoff, and its latest bounded error and attempt count stay visible in the expanded task card. Enable mcp_tasks in config.yaml, configure task_toolsets with exact raw tool names in extensions_config.json, and use a SQL database backend (sqlite or postgres). Task-enabled server connection, authentication, interceptor, timeout, or binding changes require a Gateway restart so Agent tool discovery and background calls cannot use different configuration versions. input_required is notification-only for now: DeerFlow can display the request but cannot yet submit the user's answer back to the remote task.
For deployment-level HTTP/SSE servers with task_toolsets, discovery, ordinary
tools, and background task calls share OAuth token state within one Gateway
process, including across tool-cache resets. Rotated refresh tokens stay in
memory; they are not written back to configuration or shared across processes.
After a restart, the configured refresh token must still be valid.
Notification launch and failed Agent-run deliveries use capped exponential backoff with a visible attempt count and stop after five failed attempts. When a bounded ordinary release exceeds its drain deadline, the service retains ownership until it settles. A permanently rejected target such as a deleted chat is dead-lettered immediately instead of retried forever or recreated. Cancellation endpoints return after durably recording the request; the background service owns the potentially slow remote MCP call and its retry schedule.
Notification runs keep their trusted delivery instruction separate from the framed, untrusted remote event payload. The process-started task runtime—not a hot config read—controls whether the task-management tools are exposed, so changing mcp_tasks requires a Gateway restart. When a skill's allowed-tools policy is active, list_background_tasks and cancel_background_task must be declared explicitly like other business tools.
See the MCP Server Guide for detailed instructions.
Security: pass per-request MCP credentials only through config.context.secrets;
credentials must never be placed in either run metadata surface
(metadata.auth_token or config.metadata.auth_token). See MCP credential migration and cleanup
for the supported interceptor flow and the required rotation and retained-copy
cleanup when migrating from legacy metadata credentials.
IM Channels
DeerFlow supports receiving tasks from messaging apps. Channels auto-start when configured — no public IP required for any of them.
Cancelling a channel restart discards its pending configuration reload, preserving newer runtime settings applied afterward.
DeerFlow can also expose user-owned IM channel connections in the workspace UI. When channel_connections is enabled, logged-in users can bind Telegram, Slack, Discord, Feishu/Lark, DingTalk, WeChat, WeCom, QQ, or Buzz from the sidebar / Settings > Channels. It reuses the existing outbound channels.* transports, so no public IP or provider callback URL is required. Incoming IM messages then run under the connected DeerFlow user account. See IM Channel Connections for setup and security notes.
| Channel | Transport | Difficulty |
|---|---|---|
| Telegram | Bot API (long-polling) | Easy |
| Slack | Socket Mode | Moderate |
| Feishu / Lark | WebSocket | Moderate |
| Tencent iLink (long-polling) | Moderate | |
| WeCom | WebSocket | Moderate |
| WebSocket (text-only C2C and group @mentions; four/five passive replies per source) | Moderate | |
| DingTalk | Stream Push (WebSocket) | Moderate |
| Buzz | Nostr relay (WebSocket, NIP-42) | Moderate |
Attachments saved by the shared IM ingestion pipeline or Feishu/DingTalk's embedded downloads keep distinct filenames, including when concurrent uploads choose the same name. The final filename is passed to the agent and used for sandbox sync; an existing conversation file is not overwritten.
Configuration in config.yaml:
Discord's channels.discord.allowed_guilds accepts one positive numeric guild
ID (quoted or unquoted) or a YAML list. Unset, null, [], or a blank string
allows all guilds. Invalid entries are ignored with a warning; any other
configured value yielding no valid ID denies every guild and logs an error.
allowed_channels accepts one channel ID (quoted or unquoted) or a YAML list
of IDs exempt from mention_only, within allowed guilds. An empty value gives
no exemptions, so mention_only applies everywhere when enabled.
channels:
# LangGraph-compatible Gateway API base URL (default: http://localhost:8001/api)
langgraph_url: http://localhost:8001/api
# Gateway API URL (default: http://localhost:8001)
gateway_url: http://localhost:8001
# Maximum queued or provider-reserved inbound messages (default: 1000)
inbound_queue_maxsize: 1000
# Fixed number of long-lived inbound handler workers (default: 5)
max_concurrency: 5
# Seconds to drain accepted work before cancelling active handlers (default: 3)
shutdown_grace_period_seconds: 3
# Optional: global session defaults for all mobile channels
session:
assistant_id: lead_agent # or a custom agent name; custom agents are routed via lead_agent + agent_name
config:
recursion_limit: 100
context:
thinking_enabled: true
is_plan_mode: false
subagent_enabled: false
feishu:
enabled: true
app_id: $FEISHU_APP_ID
app_secret: $FEISHU_APP_SECRET
# domain: https://open.feishu.cn # China (default)
# domain: https://open.larksuite.com # International
qq:
enabled: true
app_id: $QQ_APP_ID
client_secret: $QQ_CLIENT_SECRET
allowed_users: [] # QQ OpenIDs, not QQ account numbers
wecom:
enabled: true
bot_id: $WECOM_BOT_ID
bot_secret: $WECOM_BOT_SECRET
# Optional: extra host suffixes inbound media downloads may come from, in
# addition to the built-in qq.com family and WeCom's official COS media
# host (ww-aibot-img-1258476243..myqcloud.com); add one here if
# WeCom rotates to a new COS account or media goes through a proxy
allowed_media_hosts: []
slack:
enabled: true
bot_token: $SLACK_BOT_TOKEN # xoxb-...
app_token: $SLACK_APP_TOKEN # xapp-... (Socket Mode)
allowed_users: [] # empty = allow all
telegram:
enabled: true
bot_token: $TELEGRAM_BOT_TOKEN
# Optional: render final Markdown replies as Telegram Rich Messages.
rich_messages: false
allowed_users: [] # numeric user IDs, not @usernames; empty = allow all
wechat:
enabled: false
bot_token: $WECHAT_BOT_TOKEN
ilink_bot_id: $WECHAT_ILINK_BOT_ID
qrcode_login_enabled: true # optional: allow first-time QR bootstrap when bot_token is absent
allowed_users: [] # iLink user IDs; empty allows all; one ID is one entry
polling_timeout: 35 # timing values must be positive finite seconds
polling_retry_delay: 5
qrcode_poll_interval: 2
qrcode_poll_timeout: 180
state_dir: ./.deer-flow/wechat/state
max_inbound_image_bytes: 20971520
max_outbound_image_bytes: 20971520
max_inbound_file_bytes: 52428800
max_outbound_file_bytes: 52428800
# Inbound media downloads stream with the caps above and are restricted to
# these host suffixes (plus *.qq.com and the cdn_base_url host by default)
allowed_media_hosts: []
# Optional: per-channel / per-user session settings
session:
assistant_id: mobile-agent # custom agent names are also supported here
context:
thinking_enabled: false
users:
"123456789":
assistant_id: vip-agent
config:
recursion_limit: 150
context:
thinking_enabled: true
subagent_enabled: true
dingtalk:
enabled: true
client_id: $DINGTALK_CLIENT_ID # Client ID of your DingTalk application
client_secret: $DINGTALK_CLIENT_SECRET # Client Secret of your DingTalk application
allowed_users: [] # empty = allow all
card_template_id: "" # Optional: AI Card template ID for streaming typewriter effect
Notes:
assistant_id: lead_agentcalls the default LangGraph assistant directly.- If
assistant_idis set to a custom agent name, DeerFlow still routes throughlead_agentand injects that value asagent_name, so the custom agent's SOUL/config takes effect for IM channels. - IM channel workers call Gateway's LangGraph-compatible API internally and automatically attach process-local internal auth plus the CSRF cookie/header pair required for thread and run creation.
- Inbound work is bounded to
inbound_queue_maxsizepending messages plusmax_concurrencyactive workers. When capacity is exhausted, socket/polling providers drop new messages before sending DeerFlow's working acknowledgment and emit a rate-limited warning. Buzz leaves its replay cursor unchanged and reconnects for relay replay; GitHub webhooks return503, marking the delivery failed for manual/API redelivery. Shutdown closes admission immediately, keeps channel transports available while accepted messages drain for up toshutdown_grace_period_seconds, then cancels and awaits active handlers before closing provider resources; the Gateway's outer timeout can cancel an incomplete shutdown without detaching those resources. - Feishu/Lark now queues rapid follow-up messages per mapped DeerFlow
thread_idinstead of immediately surfacing the generic busy reply, and topic replies keep a per-message card with a compact source-message preview across queued/running/final patches. - Streaming IM channels treat backend
errorevents like transport failures and follow the same reply and retry path. The channel logs the error type and message for diagnosis.
Set the corresponding API keys in your .env file:
# Telegram
TELEGRAM_BOT_TOKEN=123456789:ABCdefGHIjklMNOpqrSTUvwxYZ
# Slack
SLACK_BOT_TOKEN=xoxb-...
SLACK_APP_TOKEN=xapp-...
# Feishu / Lark
FEISHU_APP_ID=cli_xxxx
FEISHU_APP_SECRET=your_app_secret
# WeChat iLink
WECHAT_BOT_TOKEN=your_ilink_bot_token
WECHAT_ILINK_BOT_ID=your_ilink_bot_id
# WeCom
WECOM_BOT_ID=your_bot_id
WECOM_BOT_SECRET=your_bot_secret
# DingTalk
DINGTALK_CLIENT_ID=your_client_id
DINGTALK_CLIENT_SECRET=your_client_secret
Telegram Setup
- Chat with @BotFather, send
/newbot, and copy the HTTP API token. - Set
TELEGRAM_BOT_TOKENin.envand enable the channel inconfig.yaml. - The bot accepts inbound text, photos, and documents (with or without captions). Hosted Bot API downloads are limited to 20 MB per attachment.
Slack Setup
- Create a Slack App at api.slack.com/apps → Create New App → From scratch.
- Under OAuth & Permissions, add Bot Token Scopes:
app_mentions:read,chat:write,im:history,im:read,im:write,files:write. - Enable Socket Mode → generate an App-Level Token (
xapp-…) withconnections:writescope. - Under Event Subscriptions, subscribe to bot events:
app_mention,message.im. - Set
SLACK_BOT_TOKENandSLACK_APP_TOKENin.envand enable the channel inconfig.yaml.
Feishu / Lark Setup
- Create an app on Feishu Open Platform → enable Bot capability.
- Add permissions:
im:message,im:message.p2p_msg:readonly,im:resource. - Under Events, subscribe to
im.message.receive_v1and select Long Connection mode. - Copy the App ID and App Secret. Set
FEISHU_APP_IDandFEISHU_APP_SECRETin.envand enable the channel inconfig.yaml. - The bot supports inbound text, image, and file messages. Inbound attachment downloads are limited to 20 MB per attachment.
WeChat Setup
- Enable the
wechatchannel inconfig.yaml. - Either set
WECHAT_BOT_TOKENin.env, or setqrcode_login_enabled: truefor first-time QR bootstrap. - When
bot_tokenis absent and QR bootstrap is enabled, watch backend logs for the QR content returned by iLink and complete the binding flow. - After the QR flow succeeds, DeerFlow persists the acquired token under
state_dirfor later restarts. - For Docker Compose deployments, keep
state_diron a persistent volume so theget_updates_bufcursor and saved auth state survive restarts. - Outbound images/files enforce
max_outbound_image_bytes/max_outbound_file_bytes(20 MiB / 50 MiB defaults) while reading, including files that grow after resolution. Oversize reads are rejected before encryption/upload instead of sending a truncated prefix. Non-positive limits disable the corresponding cap. allowed_userstakes iLink user IDs. Unset,null,[], or a blank string allows everyone. A single ID is one entry, not a sequence of characters, and an unquoted integer-valued number is stored as that integer's text. A scalar string containing commas or interior whitespace logs a warning but remains one literal ID; use a YAML list for multiple IDs. Invalid entries are ignored with a warning; any other configured value that yields no valid ID denies every user and logs an error./connectis still accepted before that check, and a denied sender is dropped before inbound media is downloaded.- Shutdown waits for in-flight cursor writes. On token expiry, DeerFlow persists the cursor reset and removes the saved token before completing poller cancellation.
WeCom Setup
- Create a bot on the WeCom AI Bot platform and obtain the
bot_idandbot_secret. - Enable
channels.wecominconfig.yamland fill inbot_id/bot_secret. - Set
WECOM_BOT_IDandWECOM_BOT_SECRETin.env. - Make sure backend dependencies include
wecom-aibot-python-sdk. The channel uses a WebSocket long connection and does not require a public callback URL. - The current integration supports inbound text, image, and file messages. Final images/files generated by the agent are also sent back to the WeCom conversation.
DingTalk Setup
- Create a DingTalk application in the DingTalk Developer Console and enable Robot capability.
- Set the message receiving mode to Stream Mode in the robot configuration page.
- Copy the
Client IDandClient Secret, setDINGTALK_CLIENT_IDandDINGTALK_CLIENT_SECRETin.env, and enable the channel inconfig.yaml. - (Optional) To enable streaming AI Card replies (typewriter effect), create an AI Card template on the DingTalk Card Platform, then set
card_template_idinconfig.yamlto the template ID. You also need to apply for theCard.Streaming.WriteandCard.Instance.Writepermissions.
When DeerFlow runs in Docker Compose, IM channels execute inside the gateway container. In that case, do not point channels.langgraph_url or channels.gateway_url at localhost; use container service names such as http://gateway:8001/api and http://gateway:8001, or set DEER_FLOW_CHANNELS_LANGGRAPH_URL and DEER_FLOW_CHANNELS_GATEWAY_URL.
Commands
Once a channel is connected, you can interact with DeerFlow directly from the chat:
| Command | Description |
|---|---|
/new |
Start a new conversation |
/status |
Show current thread info |
/models |
List available models |
/model [name|default] |
Show or pin the current conversation's model |
/memory |
View memory |
/agent list |
List your Custom Agents |
/agent use |
Start a new conversation with a Custom Agent |
/help |
Show help |
Messages without a command prefix are treated as regular chat — DeerFlow creates a thread and responds conversationally.
Agent selection is conversation-scoped: /agent use starts a fresh conversation and pins that Custom Agent in the thread metadata. Existing conversations never switch agents midway, the selection survives a Gateway restart, and opening the IM-created thread in the Web UI continues through the same Custom Agent.
Use /agent use lead_agent to return to the default agent in a new conversation.
Model selection is conversation-scoped too: /model pins a model to the current conversation — validated against the caller-visible model list, persisted in the thread metadata so it survives a Gateway restart, and applied from the next message without starting a new conversation. /model shows the effective model and its source, /model default clears the pin, and /models reports the pinned model.
Request Trace Correlation
Every Gateway HTTP response carries an X-Trace-Id header. The id is inherited
from an inbound X-Trace-Id when the caller sends one and generated otherwise, so
a proxy or an upstream service can pin one id across services. It needs no
configuration and cannot be turned off.
The same id stays attached to work that outlives the HTTP response: the detached
run task, any subagents it delegates to, and the background memory-update threads.
It is recorded as deerflow_trace_id on the run record (visible in the runs API),
in the thread's checkpoint metadata, and in Langfuse traces. Scheduled tasks, MCP
task notification runs, and IM channel messages start outside HTTP and mint their
own id per occurrence.
Log records carry that id only when enhanced logging is on:
logging:
enhance:
enabled: true # print trace_id into log records
format: text # or json
This is off by default because turning it on changes the log format. logging is
restart-required, so edit config.yaml and restart the Gateway. The setting
affects log output only — the id, the response header, and the run metadata are
unaffected.
deerflow_trace_id is a DeerFlow correlation id: it is not a run id, and it is not
a provider's native trace id. It is not a lookup key either — nothing resolves a
thread or a run from it; use it to correlate log lines. A deerflow_trace_id sent
in a run request's metadata or config.context is ignored and overwritten, so
the response header, the logs, and the persisted run can never disagree. To pin a
correlation id, send the X-Trace-Id header.
Gateway run history also records one terminal run.delivery receipt per run,
including zero-output and crash-recovered runs. The receipt is persisted before
the durable terminal run status during normal execution. Orphan recovery first
atomically claims an expired lease and then idempotently backfills the receipt,
so a stale recovery scan cannot overwrite a live run's detailed delivery facts.
Receipt persistence remains best-effort during an event-store outage. Runs that
fail checkpoint preflight (or are cancelled while waiting for prior
finalization) keep the existing completion-data behavior: they receive the
zero-delivery receipt but do not overwrite RunStore completion fields with an
empty snapshot.
When tool_progress.enabled is true, the same run event history also records
result-quality guard phase changes. It records loop-detection decisions and
deferred MCP tool promotions for both the lead agent and ordinary task
subagents. Promotion
events identify newly promoted deferred-tool names and whether routing metadata or
tool_search selected them, without copying the search query, routing keywords,
schemas, arguments, results, or catalog hash into the promotion event itself.
LangSmith Tracing
DeerFlow has built-in LangSmith integration for observability. When enabled, all LLM calls, agent runs, and tool executions are traced and visible in the LangSmith dashboard.
Add the following to your .env file:
LANGSMITH_TRACING=true
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
LANGSMITH_API_KEY=lsv2_pt_xxxxxxxxxxxxxxxx
LANGSMITH_PROJECT=xxx
Langfuse Tracing
DeerFlow also supports Langfuse observability for LangChain-compatible runs.
Add the following to your .env file:
LANGFUSE_TRACING=true
LANGFUSE_PUBLIC_KEY=pk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_SECRET_KEY=sk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_BASE_URL=https://cloud.langfuse.com
If you are using a self-hosted Langfuse instance, set LANGFUSE_BASE_URL to your deployment URL.
Trace correlation fields. Every agent run is annotated with Langfuse's reserved trace attributes so the Sessions and Users pages light up automatically:
session_id= LangGraphthread_id— groups every trace of the same conversationuser_id= effective user fromget_effective_user_id()(falls back todefaultin no-auth mode)trace_name= assistant id (defaults tolead-agent)tags=[env:, model:](built-in Gateway lead-agent and embedded client runs tag the selected model after default or fallback resolution; custom graph factories may only provide the requested model; the environment tag is omitted when unset)metadata.deerflow_trace_id= DeerFlow request correlation id, matchingX-Trace-Idwhen request trace correlation is enabled
These are injected into RunnableConfig.metadata at the graph invocation root for both the gateway path (runtime/runs/worker.py::run_agent) and the embedded path (client.py::DeerFlowClient.stream), so any LangChain-compatible callback can read them. Set DEER_FLOW_ENV (or ENVIRONMENT) to tag traces by deployment environment.
Monocle Tracing
DeerFlow also supports Monocle, an OpenTelemetry-based tracer for agentic applications. It records each run end-to-end: LLM calls, agent steps, and tool and MCP invocations, with their inputs, outputs, timings, and token counts.
Add the following to your .env file:
MONOCLE_TRACING=true
MONOCLE_EXPORTERS=file # file, console, okahu, s3, blob, gcs (default: file)
OKAHU_API_KEY=okh_xxxxxxxx # required only for the `okahu` exporter
Each run writes one trace file to .monocle/; open it in the Monocle VS Code extension to inspect the span timeline and token counts. Connect to Okahu, an agent-observability platform, to analyze traces across runs and run trace-based and agentic evaluations (via the okahu exporter).
Traces capture span inputs and outputs verbatim — prompts, tool arguments, and model responses — plus token usage and timings. The file exporter keeps them on local disk and never rotates or cleans them up, so prune .monocle/ periodically; the remote exporters (okahu, s3, blob, gcs) send that same data off-box, so enable only destinations you trust. Monocle is initialized once at Gateway startup: a configuration error (unknown exporter, missing OKAHU_API_KEY) is logged there and tracing stays off until the Gateway restarts.
Using Multiple Providers
LangSmith and Langfuse attach as LangChain callbacks, so you can enable both and DeerFlow reports each run to both. If an enabled provider is missing required credentials or fails to initialize, DeerFlow fails fast and names it. Monocle uses a global OpenTelemetry provider rather than a callback; Langfuse shares that provider, so all three can run together. Because both span processors sit on the same shared provider, Monocle's exporters also see Langfuse's spans when both are enabled.
For Docker deployments, tracing is disabled by default. Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in your .env to enable it.
Existing-Run Stream Actions
Existing-run SSE joins are observation-only on GET: supplying
action=interrupt|rollback returns 405. Cancellation on this stream route is
POST-only and requires the runs:cancel permission. Accordingly, the OpenAPI
contract exposes action and wait only on POST; the GET operation exposes
only its path parameters.
Personal Access Tokens
Non-interactive clients (CI pipelines, scripts, server-to-server integrations)
can call the Gateway API with a personal access token (PAT) instead of a
browser session. Create one while logged in via POST /api/v1/auth/pats — the
raw dfp_... value is shown exactly once; only its SHA-256 digest is stored —
then send it as a Bearer credential:
POST /api/threads/search
Authorization: Bearer dfp_...
Content-Type: application/json
{}
Each token runs with its owning user's identity (owner filtering and per-user
memory keep working), carries a scope set that can only narrow that user's
permissions, and is admitted only to the thread/run lifecycle routes — every
other route answers 403 to PAT callers, and a PAT never carries admin
capability. Tokens can be listed and revoked at any time; revocation is
immediate. PATs require a database backend (SQLite/PostgreSQL). Full
reference: API Reference — Personal Access Tokens.
From Deep Research to Super Agent Harness
DeerFlow started as a Deep Research framework — and the community ran with it. Since launch, developers have pushed it far beyond research: building data pipelines, generating slide decks, spinning up dashboards, automating content workflows. Things we never anticipated.
That told us something important: DeerFlow wasn't just a research tool. It was a harness — a runtime that gives agents the infrastructure to actually get work done.
So we rebuilt it from scratch.
DeerFlow 2.0 is no longer a framework you wire together. It's a super agent harness — batteries included, fully extensible. Built on LangGraph and LangChain, it ships with everything an agent needs out of the box: a filesystem, memory, skills, sandbox-aware execution, and the ability to plan and spawn sub-agents for complex, multi-step tasks.
Use it as-is. Or tear it apart and make it yours.
Core Features
Skills & Tools
Open Capability Center from the workspace sidebar to manage Plugins
(MCP servers and Lark/Feishu integration) and Skills. Both catalogs support
search; skill cards show concise descriptions with full descriptions in a detail
view. Built-in and user-created/imported skills are listed separately. The
Community tab supports importing .skill archives into My skills.
General preferences remain in Settings.
Administrators can see Could not load warnings in Capability Center > Skills
when their own custom skill's SKILL.md contains invalid YAML. Warnings show the
package-relative file and, when available, its line and column, with a quoting
hint for unquoted colon values. Fix the file on disk and select Reload skills
to refresh the list and invalidate skill prompt caches for subsequent runs;
active runs retain their existing snapshot. This covers caller-owned custom
skills only, excluding legacy/global, public, integration, and external linked
packages. These warnings cover YAML syntax errors only: missing frontmatter,
non-mapping metadata, missing/invalid name or description, invalid allowed-tools
or required-secrets declarations, non-UTF-8 files, and file-read failures
(including permission denied) are not reported. A missing warning does not
establish that a package is valid or readable. The scope note remains visible
even when there are no YAML warnings; check Gateway logs for other load failures.
Parsing remains strict and invalid skills stay unavailable to agents.
Skills are what make DeerFlow do almost anything.
A standard Agent Skill is a structured capability module — a Markdown file that defines a workflow, best practices, and references to supporting resources. DeerFlow ships with built-in skills for research, report generation, slide creation, web pages, image and video generation, and more. But the real power is extensibility: add your own skills, replace the built-in ones, or combine them into compound workflows.
Skills are loaded progressively — only when the task needs them, not all at once. This keeps the context window lean and makes DeerFlow work well even with token-sensitive models.
When deferred skill discovery is enabled, describe_skill ranks installed skills by bounded, Unicode-normalized intent-term coverage across names and descriptions. Natural multi-term requests can therefore find a relevant skill without requiring one exact phrase, while exact select: and required-name +prefix lookups remain available. Ranked searches use up to 256 characters and return up to five results; exact select: lists are not truncated and return all requested catalog matches.
A skill directory is a package boundary: once DeerFlow finds its SKILL.md, nested SKILL.md files under that package (for example evaluation fixtures) remain supporting data and are not registered as runtime skills. This applies to managed integration packs as well as public and custom skills. Namespace directories without their own SKILL.md can still group nested skills.
Discovery follows operator-managed directory symlinks, but skips links back to an ancestor directory so a cyclic namespace does not repeatedly rescan the same tree. Independent links to the same external skill tree remain supported.
Skill Markdown and bundled text resources use UTF-8. Skill-creator CLI and review utilities read and write text explicitly as UTF-8 so localized skills behave consistently across operating systems.
Custom skill history preserves Unicode line separators inside saved content and metadata, keeping those revisions readable for history inspection and rollback. Malformed JSON history records are still rejected.
The bundled github-deep-research API client also works without requests, using
Python's standard-library HTTP client. Its fallback preserves spaces, Unicode, and
reserved characters in search queries and label filters.
Users can explicitly activate an enabled skill for a single turn by starting the request with /skill-name, for example /data-analysis analyze uploads/foo.csv. DeerFlow loads that skill's SKILL.md as hidden current-turn context while leaving the base prompt limited to skill metadata. Slash activation respects disabled skills, custom-agent skill whitelists, and existing channel commands such as /new and /help.
After an answer loads skills, its toolbar includes Skills used. Hover over or click the icon to see the skills and their sources; hovering a skill name underlines it. Select a skill to inspect its SKILL.md in the resizable side panel (a drawer on mobile). The view uses snapshots captured during successful configured read-tool loads or explicit slash activation, so later edits or removal of a skill do not rewrite its history. Repeated loads appear once per run, in first-load order. Range reads and bounded snapshots are labeled as partial; older conversations without captured evidence do not show this menu. Copy returns the captured Markdown, including YAML frontmatter. Package-relative links and images remain readable references rather than navigating away from the conversation. Skill loads performed through other tools, such as shell commands, are not inferred from answer text.
An enabled skill's allowed-tools policy applies only after that skill is explicitly slash-activated or captured in the agent's active skill context after a read_file load. Merely enabling, advertising, or listing a skill in a custom agent or subagent skills allowlist does not reduce that agent's normal toolset; subagents use the same progressive discovery and activation policy as the lead agent. During a slash-activated run, that explicit skill's policy is authoritative: reading another SKILL.md may provide instructions but cannot widen the slash skill's tools. Without slash activation, policies from skills actually loaded into active context retain their union semantics. Once active, the policy filters both model-visible tool schemas and tool execution. Framework discovery tools (tool_search and describe_skill) remain available so an allowed deferred tool or installed skill can still be discovered, but discovery and promotion never grant permission to execute a business tool omitted from allowed-tools. task is not framework-exempt; a restrictive skill must list it explicitly to delegate to a subagent. Per-step policy decisions are internal runtime context and are removed from observable or persisted context copies. Registry failures and an active set with no remaining valid skill fail closed to framework-safe tools; individual stale paths are ignored only when another valid active skill remains. This is best-effort behavioral scoping, not a hard security boundary: loading skill instructions through another tool is not captured, and active-skill entries can be evicted from bounded context.
When you install .skill archives through the Gateway, DeerFlow accepts standard space-separated allowed-tools, optional frontmatter metadata, and the Claude-compatible argument-hint field instead of rejecting otherwise valid external skills. YAML lists remain supported for allowed-tools and preserve exact runtime names. Exact portable spellings such as WebFetch, WebSearch, Glob, Grep, and Read map to DeerFlow's web_fetch, web_search, glob, grep, and read_file tools; lowercase or otherwise unknown scalar names remain unchanged so custom and MCP tools keep their exact runtime spelling. Parenthesized entries such as Bash(tvly *) are tokenized as one literal entry, including spaces, quoted text, and escaped parentheses, but remain inactive because DeerFlow does not inspect tool arguments; declare bash only when the skill may use the full Bash tool.
Disabling a skill also removes it from the sandbox filesystem view, so shell commands and structured file tools follow the same enabled state. Local, Docker/AIO, hostPath provisioner, and newly created E2B sandboxes source /mnt/skills from enabled-only projections that update when public, custom, legacy, or managed integration skills are toggled, edited, created, deleted, or installed. Structured read_file calls (including line ranges and read-before-write checks) use the sandbox provider's mount mapping, so the user identity captured when the sandbox was acquired remains authoritative. Managed integration packages remain shared, while their projected filesystem visibility follows each user's enabled state. Multi-worker Gateways re-read on-disk enable state while rebuilding user projections, so a toggle handled by one worker is honored by another worker's next sandbox acquire. Existing E2B sandboxes retain their creation-time snapshot until they are recreated. PVC-backed provisioner skills keep their configured PVC snapshot/layout for now; dynamic PVC materialization is tracked separately.
# Paths inside the sandbox container
/mnt/skills/public
├── research/SKILL.md
├── report-generation/SKILL.md
├── slide-creation/SKILL.md
├── web-page/SKILL.md
└── image-generation/SKILL.md
/mnt/skills/custom
└── your-custom-skill/SKILL.md ← yours
/mnt/skills/integrations
└── lark-cli/lark-doc/SKILL.md ← managed, read-only
The built-in image-generation skill supports Gemini, MiniMax, and
OpenAI-compatible Images APIs. Select the latter with
IMAGE_GENERATION_PROVIDER=openai, then configure
IMAGE_GENERATION_API_KEY, IMAGE_GENERATION_BASE_URL, and
IMAGE_GENERATION_MODEL. For a containerized sandbox or an explicitly enabled
local host shell, expose these variables through sandbox.environment; sandbox
commands intentionally do not inherit API keys from the Gateway process.
The local provider uses the resolved configuration values unchanged, and
request-scoped secrets override operator values. Local bash output and command
errors redact credential values while leaving benign settings such as model
names and base URLs readable.
For LocalSandboxProvider, this is a managed tool-path boundary rather than host filesystem isolation. Explicit per-Agent skill policies are accepted only while host bash is disabled (the default), because a host subprocess can address canonical paths without using the provider's virtual-path mappings. Use Docker/AIO, the Kubernetes provisioner, or E2B when the filesystem boundary must remain enforceable alongside shell access.
Managed integrations install shared read-only skill packs without mixing them
into custom skills. The Lark/Feishu CLI integration is available under
Capability Center → Plugins → Lark / Feishu; an administrator installs or
upgrades the official lark-* pack once under
{DEER_FLOW_HOME}/integrations/skills/lark-cli, and every user discovers that
same pack with an independent enabled state. Each user's app configuration and
OAuth data remain isolated under
{DEER_FLOW_HOME}/users/{user_id}/integrations/lark-cli/{config,data}. These
secret directories are restricted to 0700, regular credential files to
0600, and symlinks are rejected.
After installation, users can click Connect Lark to open a browser
authorization link; no terminal authorization is required. The same UI can
request additional permission domains such as Calendar, Docs, or Drive, or a
specific OAuth scope reported by lark-cli. A cheap status refresh only
inspects the local credential tree, so the UI reports Credentials configured
(not live-verified) until an explicit browser completion performs live token
verification. The action then remains Reconnect Lark so users can replace
or extend authorization. If an agent hits missing Lark authorization during a
conversation, the managed lark-shared guidance points the user back to the
same plugin configuration with /workspace/capabilities?tab=plugins&plugin=lark.
Once configured, Change Lark app lets a user point their DeerFlow account at a different Lark/Feishu app without a reinstall — either by pasting an existing app's App ID / App Secret or by re-registering an app in the browser. Switching is per-user (it never touches another user's credentials), validates the new credentials through the official CLI's live tenant-token probe before replacing the active app, and revokes/removes the previous app's OAuth tokens. A rejected credential change does not supersede an in-progress setup or authorization flow. The previous OAuth data is cleared before the CLI stores the replacement app, so the new file-backed keychain secret remains available during reconnection. DeerFlow then immediately opens browser authorization for the newly bound app so the switch ends in a usable connection.
Installing the Lark skill pack resolves the latest official larksuite/cli
release from GitHub and downloads that version's skills at install time, so the
Gateway needs outbound internet access for that step (it falls back to a
bottom-line pinned version if the release lookup fails). The settings page shows
the installed version and, when available, the newest published version so an
admin can reinstall to upgrade. Air-gapped deployments can pre-stage the archive
and point DEER_FLOW_LARK_CLI_SKILLS_ARCHIVE at the local file. Integrity does
not depend on a pinned archive byte hash (GitHub does not guarantee stable
source-archive bytes); instead the download is restricted to the official GitHub
host, every archive member passes structural safety guards, and a content hash
of the effective installed skill tree (including DeerFlow's injected shared
guidance) is recorded so content changes are auditable across reinstalls.
When sandbox.use selects the AIO provider, the same install also downloads the
official Linux amd64 and arm64 CLI release archives, verifies their published
SHA-256 checksums, safely extracts one executable per architecture, and mounts
the resulting runtime read-only at /mnt/integrations/lark-cli/runtime. An
architecture-selecting launcher in that mount makes lark-cli available in the
sandbox PATH. Air-gapped AIO deployments can pre-stage a symlink-free runtime
tree containing bin/lark-cli plus both linux-{amd64,arm64}/lark-cli files and
set DEER_FLOW_LARK_CLI_SANDBOX_RUNTIME_DIR to that directory.
Sandbox trust boundary: the browser never receives the Lark app secret, but agent conversations run
lark-cliinside the sandbox, so the per-user credential directories are mounted into it:config(holding the long-livedappSecret) is mounted read-only, its otherwise emptyconfig/lockssubdirectory is over-mounted writable forlark-clicoordination files, anddata(refreshable OAuth tokens) is writable. The credential-bearing config and data mounts remain readable by any process the agent runs there, so code reached via prompt injection in a tool result could read them. Treat the sandbox as inside the Lark credential trust boundary until the sidecar credential-broker follow-up removes these mounts from sandbox execution.
For remote/Kubernetes deployments (the provisioner backend), the sandbox
lark-cli runtime can instead be supplied by an optional init container that
copies the binaries into a shared emptyDir — no install-time GitHub download and
no hostPath/PVC runtime mount. Publish the image under
docker/lark-cli-init and set
LARK_CLI_INIT_IMAGE on the provisioner (with the Helm chart,
provisioner.larkCliInitImage / provisioner.larkCliBrokerImage); it stays off
(legacy behavior) when unset. The Lark integration status
(GET /api/integrations/lark/status) reports sandbox_runtime_mode,
sandbox_runtime_probed, and sandbox_runtime_ready.
sandbox_runtime_probed marks whether runtime readiness was actually
evaluated; responses from older backends may omit the flag, in which case the
Settings mutation cache keeps the last probed runtime fields instead of
overwriting them with an unevaluated fallback — so the Settings UI shows
whether lark-cli will actually be present in the sandbox at chat time, rather
than a green status hiding a later command not found.
In Lark broker mode, AIO's fresh Bash runs close inherited pipe stdin. Persistent terminal commands keep their input; the shim ignores terminal stdin. Explicit pipelines, heredocs, and file redirections still supply input normally, and the shim forwards that input only after EOF. A stdin idle timeout aborts without executing the command, and broker execution logs omit argument values. See the broker image guide for timeout settings and image rebuild requirements.
If a trusted operator manages the configured skills directory through an external mount such as MinIO, NFS, or CSI, an administrator can call POST /api/skills/reload after changing files. This invalidates skill prompt caches in the handling Gateway process, waits up to the bounded refresh timeout, and then advances a durable reset marker (.extensions_config.json.skills-cache-reset.json) in the writable config directory. Every Gateway worker or Pod mounting that directory notices the marker within about a second of its next prompt build and rescans the latest files, so one call reaches every replica sharing the volume; the response reports scope: shared_config, or scope: process when no extensions config path is available. Running tasks are unchanged. Skill installs, edits, deletions, rollbacks and enable/disable toggles made through the API publish the same marker, so a change made on one replica takes effect on the others without a restart. A loader-level filesystem failure returns a generic server error and preserves the last successfully loaded process cache rather than publishing an empty catalog. Replicas with independent filesystems are not covered. Direct mount writes bypass the validation, SkillScan, and history applied by DeerFlow's install/edit APIs, so only operator-controlled systems should have write access.
Skill installs and agent-managed skill edits run through SkillScan, a native deterministic safety scanner before the LLM-based skill scanner. Phase 1 runs offline with no Semgrep/OpenGrep dependency, blocks high-confidence CRITICAL findings such as private keys or shell execution, and passes warning findings to the LLM scanner for contextual review. Code files (anything under scripts/, a script suffix such as .py, .sh, or .js, or an extensionless file starting with #!) that are not NUL-free UTF-8 text raise a warning and are still analyzed over a lossy decode, so a single stray byte cannot hide them from CRITICAL checks. The moderation adapter normalizes both plain-text model responses and LangChain Responses API text blocks before parsing the required JSON decision. Python instance-client exfiltration checks follow a minimal same-scope evidence chain: a simple name bound to a known client constructor, optional name-to-name aliases, and an actual outbound method or context-manager use supported by that constructor. Constructor roots must be proven imports; bare canonical-looking names are not inferred as modules. Nested scopes do not inherit client handles and inherit only constructor import aliases that are never rebound in the enclosing scope. Comprehensions, walrus-bearing statements, annotations, complex binding targets, unsupported operations, and ambiguous branch flows produce no finding from this signal; skipped constructs conservatively invalidate every name they may bind so
