Claw Code Agent logo

Claw Code Agent

A Python reimplementation of the Claude Code agent architecture โ€” local models, full control, zero dependencies.

Python 3.10+ GitHub vLLM Qwen3-Coder Zero Dependencies Alpha License

--- ## ๐Ÿ“ข What's New > **April 2026 โ€” Major Update** | | Feature | Details | |---|---------|---------| | ๐Ÿ†• | **Interactive Chat Mode** | New `agent-chat` command โ€” multi-turn REPL with `/exit` to quit | | ๐Ÿ†• | **Streaming Output** | Token-by-token streaming with `--stream` flag | | ๐Ÿ†• | **Plugin Runtime** | Full manifest-based plugin system โ€” hooks, tool aliases, virtual tools, tool blocking | | ๐Ÿ†• | **Nested Agent Delegation** | Delegate subtasks to child agents with dependency-aware topological batching | | ๐Ÿ†• | **Agent Manager** | Lineage tracking, group membership, batch summaries for nested agents | | ๐Ÿ†• | **Cost Tracking & Budgets** | Token budgets, cost budgets, tool-call limits, model-call limits, session-turn limits | | ๐Ÿ†• | **Structured Output** | JSON schema response mode with `--response-schema-file` | | ๐Ÿ†• | **Context Compaction** | Auto-snip, auto-compact, and reactive compaction on prompt-too-long errors | | ๐Ÿ†• | **File History Replay** | Journaling of file edits with snapshot IDs, replay summaries on session resume | | ๐Ÿ†• | **Truncation Continuation** | Automatic continuation when model response is cut off (`finish_reason=length`) | | ๐Ÿ†• | **Ollama Support** | Works out of the box with Ollama's OpenAI-compatible API | | ๐Ÿ†• | **LiteLLM Proxy Support** | Route through LiteLLM Proxy to any provider | | ๐Ÿ†• | **OpenRouter Support** | Cloud API gateway โ€” access OpenAI, Anthropic, Google models via one endpoint | | ๐Ÿ†• | **Query Engine** | Runtime event counters, transcript summaries, orchestration reports | | ๐Ÿ†• | **Remote Runtime** | Manifest-backed local remote profiles, connect/disconnect state, and remote CLI/slash flows | | ๐Ÿ†• | **Daemon Commands** | Local `daemon start/ps/logs/attach/kill` wrapper over background agent sessions | | ๐Ÿ†• | **Testing Guide** | Comprehensive [TESTING_GUIDE.md](TESTING_GUIDE.md) with commands for every feature | | ๐Ÿ†• | **Parity Checklist** | Full [PARITY_CHECKLIST.md](PARITY_CHECKLIST.md) tracking implementation status vs npm source | --- ## ๐Ÿ“– About This repository reimplements the [Claude Code](https://docs.anthropic.com/en/docs/claude-code) npm agent architecture **entirely in Python**, designed to run with **local open-source models** via an OpenAI-compatible API server. Built on the public porting workspace from [instructkr/claw-code](https://github.com/instructkr/claw-code), the active development lives at [HarnessLab/claw-code-agent](https://github.com/HarnessLab/claw-code-agent). > **Goal:** Not to ship the original npm source, but to reimplement the full agent flow in Python โ€” prompt assembly, context building, slash commands, tool calling, session persistence, and local model execution. > > **Zero external dependencies** โ€” just Python's standard library.

Claw Code Agent demo

--- ## โœจ Key Features | Feature | Description | |---------|-------------| | ๐Ÿค– **Agent Loop** | Full agentic coding loop with tool calling and iterative reasoning | | ๐Ÿ’ฌ **Interactive Chat** | Multi-turn REPL via `agent-chat` with session continuity | | ๐Ÿงฐ **Core Tools** | File read / write / edit, glob search, grep search, shell execution | | ๐Ÿ”Œ **Plugin Runtime** | Manifest-based plugins with hooks, aliases, virtual tools, and tool blocking | | ๐Ÿช† **Nested Delegation** | Delegate subtasks to child agents with dependency-aware topological batching | | ๐Ÿ“ก **Streaming** | Token-by-token streaming output with `--stream` | | ๐Ÿ’ฌ **Slash Commands** | Local commands: `/help`, `/context`, `/tools`, `/memory`, `/status`, `/model`, and more | | ๐ŸŒ **Remote Runtime** | Manifest-backed remote profiles with local `remote-mode`, `ssh-mode`, `teleport-mode`, and connect/disconnect state | | ๐Ÿง  **Context Engine** | Automatic context building with CLAUDE.md discovery, compaction, and snipping | | ๐Ÿ”„ **Session Persistence** | Save and resume agent sessions with file-history replay | | ๐Ÿ’ฐ **Cost & Budget Control** | Token budgets, cost limits, tool-call caps, model-call caps | | ๐Ÿ“‹ **Structured Output** | JSON schema response mode for programmatic use | | ๐Ÿ” **Permission System** | Granular control: `--allow-write`, `--allow-shell`, `--unsafe` | | ๐Ÿ—๏ธ **OpenAI-Compatible** | Works with vLLM, Ollama, LiteLLM Proxy, OpenRouter โ€” any OpenAI-compatible API | | ๐Ÿ‰ **Qwen3-Coder** | First-class support for `Qwen3-Coder-30B-A3B-Instruct` via vLLM | | ๐Ÿ“ฆ **Zero Dependencies** | Pure Python standard library โ€” nothing to install | --- ## ๐Ÿ“‹ Roadmap ### ๐Ÿ“š Documentation | Document | Description | |----------|-------------| | [TESTING_GUIDE.md](TESTING_GUIDE.md) | Step-by-step commands to verify every feature | | [PARITY_CHECKLIST.md](PARITY_CHECKLIST.md) | Full implementation status vs the npm source | ### โœ… Done - [x] Python CLI agent loop - [x] Interactive chat mode (`agent-chat`) with multi-turn REPL - [x] OpenAI-compatible local model backend - [x] Qwen3-Coder support through vLLM with `qwen3_xml` tool parser - [x] Ollama, LiteLLM Proxy, and OpenRouter backends - [x] Core tools: `list_dir`, `read_file`, `write_file`, `edit_file`, `glob_search`, `grep_search`, `bash` - [x] Context building and `/context`-style usage reporting - [x] Slash commands: `/help`, `/context`, `/context-raw`, `/prompt`, `/permissions`, `/model`, `/tools`, `/memory`, `/status`, `/clear` - [x] Session persistence and `agent-resume` flow - [x] Permission system (read-only, write, shell, unsafe tiers) - [x] Streaming token-by-token assistant output - [x] Truncated-response continuation flow - [x] Auto-snip and auto-compact context reduction - [x] Reactive compaction retry on prompt-too-long errors - [x] Cost tracking and usage budget enforcement - [x] Token, tool-call, model-call, and session-turn budgets - [x] Structured output / JSON schema response mode - [x] File history journaling with snapshot IDs and replay summaries - [x] Nested agent delegation with dependency-aware topological batching - [x] Agent manager with lineage tracking and group membership - [x] Local daemon-style background command family - [x] Local remote runtime: manifest discovery, profile listing, connect/disconnect persistence, and CLI/slash flows - [x] Plugin runtime: manifest discovery, hooks, aliases, virtual tools, tool blocking - [x] Plugin lifecycle hooks: resume, persist, delegate phases - [x] Plugin session-state persistence and resume restoration - [x] Query engine facade driving the real Python runtime - [x] Compaction metadata with lineage IDs and revision summaries - [x] Unit tests for the Python runtime - [x] `pyproject.toml` packaging with `setuptools` ### ๐Ÿ”ฒ In Progress - [ ] Full MCP server support - [ ] Full slash-command parity with npm runtime - [ ] Full interactive REPL / TUI behavior - [ ] Exact tokenizer-accurate context accounting - [ ] Hooks system parity - [ ] Real remote transport/runtime parity beyond the current local remote-profile runtime - [ ] Voice and VIM modes - [ ] Editor and platform integrations - [ ] Background and team features --- ## ๐Ÿ—๏ธ Architecture ```text claw-code/ โ”œโ”€โ”€ README.md โ”œโ”€โ”€ TESTING_GUIDE.md # How to test every feature โ”œโ”€โ”€ PARITY_CHECKLIST.md # Implementation status vs npm source โ”œโ”€โ”€ pyproject.toml โ”œโ”€โ”€ .gitignore โ”œโ”€โ”€ images/ โ”‚ โ””โ”€โ”€ logo.png โ”œโ”€โ”€ src/ # Python implementation (75+ modules) โ”‚ โ”œโ”€โ”€ main.py # CLI entry point & argument parsing โ”‚ โ”œโ”€โ”€ agent_runtime.py # Core agent loop (LocalCodingAgent) โ”‚ โ”œโ”€โ”€ agent_tools.py # Tool definitions & execution engine โ”‚ โ”œโ”€โ”€ agent_prompting.py # System prompt assembly โ”‚ โ”œโ”€โ”€ agent_context.py # Context building & CLAUDE.md discovery โ”‚ โ”œโ”€โ”€ agent_context_usage.py # Context usage estimation & reporting โ”‚ โ”œโ”€โ”€ agent_session.py # Session state management โ”‚ โ”œโ”€โ”€ agent_slash_commands.py # Local slash command processing โ”‚ โ”œโ”€โ”€ agent_manager.py # Nested agent lineage & group tracking โ”‚ โ”œโ”€โ”€ agent_types.py # Shared dataclasses & type definitions โ”‚ โ”œโ”€โ”€ openai_compat.py # OpenAI-compatible API client (streaming) โ”‚ โ”œโ”€โ”€ plugin_runtime.py # Plugin manifest, hooks, aliases, virtual tools โ”‚ โ”œโ”€โ”€ agent_plugin_cache.py # Plugin discovery & prompt injection cache โ”‚ โ”œโ”€โ”€ session_store.py # Session serialization & persistence โ”‚ โ”œโ”€โ”€ transcript.py # Transcript block export & mutation tracking โ”‚ โ”œโ”€โ”€ query_engine.py # Query engine facade & runtime orchestration โ”‚ โ”œโ”€โ”€ remote_runtime.py # Local remote profiles, connect/disconnect state, remote CLI support โ”‚ โ”œโ”€โ”€ account_runtime.py # Local account profiles, login/logout state, account CLI support โ”‚ โ”œโ”€โ”€ config_runtime.py # Local workspace config/settings discovery and mutation โ”‚ โ”œโ”€โ”€ permissions.py # Tool permission filtering โ”‚ โ”œโ”€โ”€ cost_tracker.py # Cost & budget enforcement โ”‚ โ”œโ”€โ”€ tools.py # Mirrored tool inventory โ”‚ โ”œโ”€โ”€ commands.py # Mirrored command inventory โ”‚ โ”œโ”€โ”€ plugins/ # Plugin subsystem โ”‚ โ”œโ”€โ”€ hooks/ # Hook system (WIP) โ”‚ โ”œโ”€โ”€ remote/ # Remote runtime modes (WIP) โ”‚ โ”œโ”€โ”€ voice/ # Voice mode (WIP) โ”‚ โ””โ”€โ”€ vim/ # VIM mode (WIP) โ””โ”€โ”€ tests/ # Unit tests โ”œโ”€โ”€ test_agent_runtime.py โ”œโ”€โ”€ test_agent_context.py โ”œโ”€โ”€ test_agent_context_usage.py โ”œโ”€โ”€ test_agent_prompting.py โ”œโ”€โ”€ test_agent_slash_commands.py โ”œโ”€โ”€ test_main.py โ”œโ”€โ”€ test_query_engine_runtime.py โ””โ”€โ”€ test_porting_workspace.py ``` --- ## ๐Ÿ“ฆ Requirements | Requirement | Details | |-------------|---------| | ๐Ÿ Python | `3.10` or higher | | ๐Ÿ“š Dependencies | **None** โ€” pure Python standard library | | ๐Ÿ–ฅ๏ธ Model Server | `vLLM`, `Ollama`, `LiteLLM Proxy`, or `OpenRouter`, with tool calling support | | ๐Ÿง  Model | [`Qwen/Qwen3-Coder-30B-A3B-Instruct`](https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct) (recommended) | --- ## ๐Ÿš€ Quick Start ### 1. Start vLLM with Qwen3-Coder vLLM must be started with automatic tool choice enabled. Use the `qwen3_xml` parser for Qwen3-Coder tool calling: ```bash python -m vllm.entrypoints.openai.api_server \ --model Qwen/Qwen3-Coder-30B-A3B-Instruct \ --host 127.0.0.1 \ --port 8000 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_xml ``` Verify the server is running: ```bash curl http://127.0.0.1:8000/v1/models ``` > ๐Ÿ“š **References:** [vLLM Tool Calling Docs](https://docs.vllm.ai/en/v0.13.0/features/tool_calling/) ยท [OpenAI-Compatible Server](https://docs.vllm.ai/en/v0.13.0/serving/openai_compatible_server.html) ### Optional: Use Ollama Instead of vLLM `claw-code-agent` can also work with Ollama because the runtime targets an OpenAI-compatible API. Use a model that supports tool calling well. Example: ```bash ollama serve ollama pull qwen3 ``` Then configure: ```bash export OPENAI_BASE_URL=http://127.0.0.1:11434/v1 export OPENAI_API_KEY=ollama export OPENAI_MODEL=qwen3 ``` Notes: - prefer tool-capable models such as `qwen3` - plain chat-only models are not enough for full agent behavior - Ollama does not use the `vLLM` parser flags shown above > ๐Ÿ“š **References:** [Ollama OpenAI Compatibility](https://docs.ollama.com/api/openai-compatibility) ยท [Ollama Tool Calling](https://docs.ollama.com/capabilities/tool-calling) ### Optional: Use LiteLLM Proxy `claw-code-agent` can also work through LiteLLM Proxy because the runtime targets an OpenAI-compatible chat completions API. The routed model still needs to support tool calling for full agent behavior. Quick start example: ```bash pip install 'litellm[proxy]' litellm --model ollama/qwen3 ``` LiteLLM Proxy runs on port `4000` by default. Then configure: ```bash export OPENAI_BASE_URL=http://127.0.0.1:4000 export OPENAI_API_KEY=anything export OPENAI_MODEL=ollama/qwen3 ``` Notes: - LiteLLM Proxy gives you an OpenAI-style gateway in front of many providers - tool use still depends on the underlying routed model and provider behavior - if you configure a LiteLLM master key, use that instead of `anything` > ๐Ÿ“š **References:** [LiteLLM Docs](https://docs.litellm.ai/) ยท [LiteLLM Proxy Quick Start](https://docs.litellm.ai/) ### Optional: Use OpenRouter `claw-code-agent` can also work with [OpenRouter](https://openrouter.ai/), a cloud API gateway that provides access to models from OpenAI, Anthropic, Google, Meta, and others through a single OpenAI-compatible endpoint. No local model server required. Configure: ```bash export OPENAI_BASE_URL=https://openrouter.ai/api/v1 export OPENAI_API_KEY=sk-or-v1-your-key-here export OPENAI_MODEL=openai/gpt-4o-mini ``` Notes: - sign up at [openrouter.ai](https://openrouter.ai/) and create an API key under [Keys](https://openrouter.ai/keys) - model names use the `provider/model` format (e.g. `anthropic/claude-sonnet-4`, `openai/gpt-4o`, `google/gemini-2.5-pro`) - tool calling support varies by model โ€” check the [model list](https://openrouter.ai/models) for capabilities - this sends your conversation (including file contents and shell output) to OpenRouter and the upstream provider โ€” do not use with repos containing secrets or sensitive data > ๐Ÿ“š **References:** [OpenRouter Docs](https://openrouter.ai/docs) ยท [Supported Models](https://openrouter.ai/models) ยท [API Keys](https://openrouter.ai/keys) ### 2. Configure Environment ```bash export OPENAI_BASE_URL=http://127.0.0.1:8000/v1 export OPENAI_API_KEY=local-token export OPENAI_MODEL=Qwen/Qwen3-Coder-30B-A3B-Instruct ``` ### Use Another Model With vLLM If you want to try another model, keep the same `vLLM` server setup and change the `--model` value when you launch `vLLM`. Example: ```bash python -m vllm.entrypoints.openai.api_server \ --model your-model-name \ --host 127.0.0.1 \ --port 8000 \ --enable-auto-tool-choice \ --tool-call-parser your_parser ``` Then update: ```bash export OPENAI_MODEL=your-model-name ``` Notes: - the documented path in this repository is `vLLM` - the model must support tool calling well enough for agent use - some model families require a different `--tool-call-parser` - slash commands such as `/help`, `/context`, and `/tools` are local and do not require the model server ### 3. Run the Agent ```bash # Read-only question python3 -m src.main agent \ "Read src/agent_runtime.py and summarize how the loop works." \ --cwd . # Write-enabled task python3 -m src.main agent \ "Create TEST_QWEN_AGENT.md with one line: test ok" \ --cwd . --allow-write # Shell-enabled task python3 -m src.main agent \ "Run pwd and ls src, then summarize the result." \ --cwd . --allow-shell # Interactive chat mode python3 -m src.main agent-chat --cwd . # Streaming output python3 -m src.main agent \ "Explain the current architecture." \ --cwd . --stream ``` --- ## ๐Ÿ› ๏ธ Usage ### Agent Commands | Command | Description | |---------|-------------| | `agent ` | Run the agent with a prompt | | `agent-chat [prompt]` | Start interactive multi-turn chat mode | | `agent-prompt` | Show the assembled system prompt | | `agent-context` | Show estimated context usage | | `agent-context-raw` | Show the raw context snapshot | | `agent-resume ` | Resume a saved session | ### CLI Flags | Flag | Description | |------|-------------| | `--cwd ` | Set the workspace directory | | `--model ` | Override the model name | | `--base-url ` | Override the API base URL | | `--allow-write` | Allow the agent to modify files | | `--allow-shell` | Allow the agent to execute shell commands | | `--unsafe` | Allow destructive shell operations | | `--stream` | Enable token-by-token streaming output | | `--show-transcript` | Print the full message transcript | | `--system-prompt ` | Set a custom system prompt | | `--append-system-prompt ` | Append to the system prompt | | `--add-dir ` | Add extra directories to context | ### Budget & Limit Flags | Flag | Description | |------|-------------| | `--max-total-tokens ` | Total token budget | | `--max-input-tokens ` | Input token budget | | `--max-output-tokens ` | Output token budget | | `--max-reasoning-tokens ` | Reasoning token budget | | `--max-budget-usd ` | Maximum cost in USD | | `--max-tool-calls ` | Maximum tool calls per run | | `--max-delegated-tasks ` | Maximum delegated subtasks | | `--max-model-calls ` | Maximum model API calls | | `--max-session-turns ` | Maximum session turns | | `--input-cost-per-million ` | Input token pricing | | `--output-cost-per-million ` | Output token pricing | ### Context Control Flags | Flag | Description | |------|-------------| | `--auto-snip-threshold ` | Auto-snip older messages at this token count | | `--auto-compact-threshold ` | Auto-compact at this token count | | `--compact-preserve-messages ` | Messages to preserve during compaction | | `--disable-claude-md` | Disable CLAUDE.md discovery | ### Structured Output Flags | Flag | Description | |------|-------------| | `--response-schema-file ` | JSON schema file for structured output | | `--response-schema-name ` | Schema name identifier | | `--response-schema-strict` | Enforce strict schema validation | ### Slash Commands These are handled **locally** before the model loop: | Command | Aliases | Description | |---------|---------|-------------| | `/help` | `/commands` | Show built-in slash commands | | `/context` | `/usage` | Show estimated session context usage | | `/context-raw` | `/env` | Show raw environment & context snapshot | | `/prompt` | `/system-prompt` | Render the effective system prompt | | `/permissions` | โ€” | Show active tool permission mode | | `/model` | โ€” | Show or update the active model | | `/tools` | โ€” | List registered tools with permission status | | `/memory` | โ€” | Show loaded CLAUDE.md memory bundle | | `/status` | `/session` | Show runtime/session status summary | | `/clear` | โ€” | Clear ephemeral runtime state | ```bash python3 -m src.main agent "/help" python3 -m src.main agent "/context" --cwd . python3 -m src.main agent "/tools" --cwd . python3 -m src.main agent "/status" --cwd . ``` ### Utility Commands ```bash python3 -m src.main summary # Workspace summary python3 -m src.main manifest # Workspace manifest python3 -m src.main commands --limit 10 # Command inventory python3 -m src.main tools --limit 10 # Tool inventory ``` --- ## ๐Ÿ”ง Built-in Tools The agent has access to 7 core tools: | Tool | Description | Permission | |------|-------------|------------| | `list_dir` | List files and directories | ๐ŸŸข Always | | `read_file` | Read file contents (with line ranges) | ๐ŸŸข Always | | `write_file` | Write or create files | ๐ŸŸก `--allow-write` | | `edit_file` | Edit files via exact string matching | ๐ŸŸก `--allow-write` | | `glob_search` | Find files by glob pattern | ๐ŸŸข Always | | `grep_search` | Search file contents by regex | ๐ŸŸข Always | | `bash` | Execute shell commands | ๐Ÿ”ด `--allow-shell` | --- ## ๐Ÿ”Œ Plugin System Claw Code Agent supports a **manifest-based plugin runtime**. Drop a `plugin.json` in a `plugins/` subdirectory: ```json { "name": "my-plugin", "hooks": { "beforePrompt": "Inject guidance into the system prompt.", "afterTurn": "Run after each agent turn.", "onResume": "Reapply state on session resume.", "beforePersist": "Save state before session is saved.", "beforeDelegate": "Inject guidance before child agents.", "afterDelegate": "Process child agent results." }, "toolAliases": [ { "name": "my_read", "baseTool": "read_file", "description": "Custom read alias." } ], "virtualTools": [ { "name": "my_tool", "description": "A virtual tool.", "responseTemplate": "result: {input}" } ] } ``` > See [TESTING_GUIDE.md](TESTING_GUIDE.md) **Section 13** for full plugin testing commands. --- ## ๐Ÿช† Nested Agent Delegation The agent can delegate subtasks to child agents with full context carryover: ```bash python3 -m src.main agent \ "Delegate a subtask to inspect src/agent_runtime.py and return a summary." \ --cwd . --show-transcript ``` Features: - Sequential and parallel subtask execution - Dependency-aware topological batching - Child-session save and resume - Agent manager lineage tracking > See [TESTING_GUIDE.md](TESTING_GUIDE.md) **Section 12** for delegation testing commands. --- ## ๐Ÿ”„ Session Persistence Each `agent` run automatically saves a resumable session: ```text session_id=4f2c8c6f9c0e4d7c9c7b1b2a3d4e5f67 session_path=.port_sessions/agent/4f2c8c6f... ``` Resume a previous session: ```bash python3 -m src.main agent-resume \ 4f2c8c6f9c0e4d7c9c7b1b2a3d4e5f67 \ "Continue the previous task and finish the missing parts." ``` Resume directly into interactive chat: ```bash python3 -m src.main agent-chat \ --resume-session-id \ --cwd . ``` Inspect saved sessions: ```bash ls -lt .port_sessions/agent ``` > **Note:** Run `agent-resume` from the same `claw-code/` directory where the session was created. A resumed session continues from the saved transcript, not from scratch. --- ## ๐Ÿงช Testing Run the full test suite: ```bash python3 -m unittest discover -s tests -v ``` Smoke tests: ```bash python3 -m src.main agent "/help" python3 -m src.main agent-context --cwd . python3 -m src.main agent \ "Read src/agent_session.py and summarize the message flow." \ --cwd . ``` > ๐Ÿ“š **Full testing guide:** See [TESTING_GUIDE.md](TESTING_GUIDE.md) for step-by-step commands covering all 16 feature areas. --- ## ๐Ÿ” Permission Model Claw Code Agent uses a **tiered permission system** to keep the agent safe by default: | Tier | Capability | Flag Required | |------|-----------|---------------| | **Read-only** | List, read, glob, grep | None (default) | | **Write** | + file creation and editing | `--allow-write` | | **Shell** | + shell command execution | `--allow-shell` | | **Unsafe** | + destructive shell operations | `--unsafe` | --- ## ๐Ÿ”Ž Parity Status The full implementation checklist tracking parity against the npm `src` lives in [PARITY_CHECKLIST.md](PARITY_CHECKLIST.md). It covers: core runtime, CLI modes, prompt assembly, context/memory, slash commands, tools, permissions, plugins, MCP, REPL/TUI, remote features, editor integrations, and internal subsystems. --- ## โš ๏ธ Disclaimer - This repository is a **Python reimplementation** inspired by the Claude Code npm architecture. - It does **not** ship the original npm source. - It is **not** affiliated with or endorsed by Anthropic. ---

Built with ๐Ÿ Python ยท Powered by ๐Ÿ‰ HarnessLab Team.