Files
zk-data-agent/TESTING_GUIDE.md
T
Abdelrahman Abdallah 783145fe6a add mcp and online search
2026-04-05 02:35:49 +02:00

33 KiB

Testing Guide

This guide gives concrete commands you can run to verify the current Python implementation feature by feature.

All commands below assume you are inside:

cd /path/to/claw-code-agent

1. Environment Setup

1.1 Start vLLM with Qwen3-Coder

python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen3-Coder-30B-A3B-Instruct \
  --host 127.0.0.1 \
  --port 8000 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml

Verify the server:

curl http://127.0.0.1:8000/v1/models

Set the runtime environment:

export OPENAI_BASE_URL=http://127.0.0.1:8000/v1
export OPENAI_API_KEY=local-token
export OPENAI_MODEL=Qwen/Qwen3-Coder-30B-A3B-Instruct

1.2 Run the unit test suite

python3 -m unittest discover -s tests -v

2. Installation And Basic Usage

2.1 Editable install

pip install -e .

2.2 Show CLI help

python3 -m src.main --help
python3 -m src.main agent --help
python3 -m src.main agent-bg --help
python3 -m src.main agent-ps --help
python3 -m src.main agent-chat --help
python3 -m src.main agent-resume --help

2.3 Run the packaged entrypoint

claw-code-agent agent "/help"

3. Slash Commands

These are handled locally and do not require the model to answer.

python3 -m src.main agent "/help"
python3 -m src.main agent "/commands"
python3 -m src.main agent "/context" --cwd ..
python3 -m src.main agent "/context-raw" --cwd ..
python3 -m src.main agent "/search" --cwd ..
python3 -m src.main agent "/remote" --cwd ..
python3 -m src.main agent "/remotes" --cwd ..
python3 -m src.main agent "/plan" --cwd ..
python3 -m src.main agent "/prompt" --cwd ..
python3 -m src.main agent "/permissions" --cwd ..
python3 -m src.main agent "/hooks" --cwd ..
python3 -m src.main agent "/policy" --cwd ..
python3 -m src.main agent "/trust" --cwd ..
python3 -m src.main agent "/task-next" --cwd ..
python3 -m src.main agent "/tools" --cwd ..
python3 -m src.main agent "/memory" --cwd ..
python3 -m src.main agent "/status" --cwd ..
python3 -m src.main agent "/clear" --cwd ..

4. Context And Prompt Inspection

4.1 System prompt rendering

python3 -m src.main agent-prompt --cwd ..

4.2 Context usage accounting

python3 -m src.main agent-context --cwd ..

4.3 Raw context snapshot

python3 -m src.main agent-context-raw --cwd ..

4.4 Additional working directories

python3 -m src.main agent-context --cwd .. --add-dir /path/to/directory

4.5 Disable CLAUDE.md discovery

python3 -m src.main agent-context --cwd .. --disable-claude-md

4.6 Hook/policy context and trust inspection

Create a local policy file:

mkdir -p ./test_cases
cat > ./test_cases/.claw-policy.json <<'EOF'
{
  "trusted": false,
  "managedSettings": {
    "reviewMode": "strict"
  },
  "safeEnv": ["HOOK_SAFE_TOKEN"],
  "hooks": {
    "beforePrompt": ["Respect workspace policy before acting."],
    "afterTurn": ["Persist the policy decision after each turn."],
    "beforeTool": {
      "read_file": ["Validate the path before reading."]
    }
  }
}
EOF
export HOOK_SAFE_TOKEN=demo-secret

Inspect the runtime view:

python3 -m src.main agent "/hooks" --cwd ./test_cases
python3 -m src.main agent "/trust" --cwd ./test_cases
python3 -m src.main agent "/permissions" --cwd ./test_cases
python3 -m src.main agent "/tools" --cwd ./test_cases
python3 -m src.main agent-context-raw --cwd ./test_cases
python3 -m src.main agent-prompt --cwd ./test_cases

4.7 Safe environment values in shell tools

python3 -m src.main agent \
  "Run bash and print HOOK_SAFE_TOKEN, then explain where it came from." \
  --cwd ./test_cases \
  --allow-shell \
  --show-transcript

4.8 Plan runtime context and prompt inspection

python3 -m src.main agent \
  "Use update_plan to store a two-step plan for inspecting and editing the workspace." \
  --cwd ./test_cases \
  --allow-write \
  --show-transcript

python3 -m src.main agent "/plan" --cwd ./test_cases
python3 -m src.main agent "/tasks" --cwd ./test_cases
python3 -m src.main agent-context-raw --cwd ./test_cases
python3 -m src.main agent-prompt --cwd ./test_cases

4.9 Remote runtime context and prompt inspection

Create a local remote manifest:

cat > ./test_cases/.claw-remote.json <<'EOF'
{
  "profiles": [
    {
      "name": "staging",
      "mode": "ssh",
      "target": "dev@staging",
      "workspaceCwd": "/srv/app",
      "sessionUrl": "wss://remote/session"
    },
    {
      "name": "preview",
      "mode": "deep-link",
      "target": "preview://session"
    }
  ]
}
EOF

Inspect the runtime view:

python3 -m src.main agent "/remote" --cwd ./test_cases
python3 -m src.main agent "/remotes" --cwd ./test_cases
python3 -m src.main agent-context-raw --cwd ./test_cases
python3 -m src.main agent-prompt --cwd ./test_cases

4.10 Search runtime context and prompt inspection

Create a local search manifest:

cat > ./test_cases/.claw-search.json <<'EOF'
{
  "providers": [
    {
      "name": "local-search",
      "provider": "searxng",
      "baseUrl": "http://127.0.0.1:8080"
    }
  ]
}
EOF

Inspect the runtime view:

python3 -m src.main agent "/search" --cwd ./test_cases
python3 -m src.main agent "/search providers" --cwd ./test_cases
python3 -m src.main agent-context-raw --cwd ./test_cases
python3 -m src.main agent-prompt --cwd ./test_cases

4.11 Tokenizer-aware context accounting

Inspect which token counter backend the runtime is using:

python3 -m src.main agent "/status" --cwd ./test_cases
python3 -m src.main agent-context --cwd ./test_cases

Force a local tokenizer path or model override for context accounting:

export CLAW_CODE_TOKENIZER_PATH=/path/to/local/tokenizer
# or
export CLAW_CODE_TOKENIZER_MODEL=Qwen/Qwen3-Coder-30B-A3B-Instruct

python3 -m src.main agent "/status" --cwd ./test_cases
python3 -m src.main agent-context --cwd ./test_cases

If no tokenizer backend is available, the runtime will fall back to a heuristic counter and /status will show that.

5. Core Agent Loop

5.1 Read-only run

python3 -m src.main agent \
  "Read claw-code-agent/src/agent_runtime.py and summarize how the loop works." \
  --cwd ..

5.2 Show transcript output

python3 -m src.main agent \
  "Read claw-code-agent/src/agent_session.py and summarize the session model." \
  --cwd .. \
  --show-transcript

5.3 Streaming model responses

python3 -m src.main agent \
  "Inspect the current repository and summarize the architecture." \
  --cwd .. \
  --stream \
  --show-transcript

5.4 Interactive chat loop

python3 -m src.main agent-chat --cwd ..

Optional first prompt:

python3 -m src.main agent-chat \
  "Inspect the repository and tell me where the runtime loop lives." \
  --cwd ..

6.1 Configure a provider with a local manifest

cat > ./test_cases/.claw-search.json <<'EOF'
{
  "providers": [
    {
      "name": "local-search",
      "provider": "searxng",
      "baseUrl": "http://127.0.0.1:8080"
    },
    {
      "name": "backup-search",
      "provider": "tavily",
      "apiKeyEnv": "TAVILY_API_KEY"
    }
  ]
}
EOF

6.2 Configure a provider from environment variables

Use one of these:

export SEARXNG_BASE_URL=http://127.0.0.1:8080
export BRAVE_SEARCH_API_KEY=your-brave-key
export TAVILY_API_KEY=your-tavily-key

6.3 Inspect providers from the CLI

python3 -m src.main search-status --cwd ./test_cases
python3 -m src.main search-providers --cwd ./test_cases
python3 -m src.main search-status --cwd ./test_cases --provider local-search
python3 -m src.main search-activate local-search --cwd ./test_cases

6.4 Run a real web search from the CLI

python3 -m src.main search \
  "python argparse mutually exclusive group" \
  --cwd ./test_cases \
  --provider local-search \
  --max-results 5

Limit results to specific domains:

python3 -m src.main search \
  "OpenAI Responses API" \
  --cwd ./test_cases \
  --domain openai.com \
  --domain platform.openai.com

6.5 Run a real web search through slash commands

python3 -m src.main agent "/search" --cwd ./test_cases
python3 -m src.main agent "/search providers" --cwd ./test_cases
python3 -m src.main agent "/search provider local-search" --cwd ./test_cases
python3 -m src.main agent "/search use local-search" --cwd ./test_cases
python3 -m src.main agent "/search python unittest mock patch examples" --cwd ./test_cases

6.6 Run a real web search through the model tool loop

python3 -m src.main agent \
  "Use web_search to find Python unittest mocking references, then summarize the top results." \
  --cwd ./test_cases \
  --show-transcript

Inside chat mode:

  • type normal prompts to continue the same session
  • use /exit or /quit to leave

5.5 Resume directly into chat mode

python3 -m src.main agent-chat \
  --resume-session-id <session-id> \
  --cwd ..

6. Tool Execution

6.1 Read files

python3 -m src.main agent \
  "Read claw-code-agent/src/agent_tools.py and summarize each tool." \
  --cwd ..

6.2 Write files

python3 -m src.main agent \
  "Create TEST_WRITE.md in the current directory with one line: write test ok" \
  --cwd ./test_cases \
  --allow-write

6.3 Edit files

python3 -m src.main agent \
  "Create demo.txt with 'hello world', then replace 'world' with 'agent'." \
  --cwd ./test_cases \
  --allow-write

6.4 Glob and grep

python3 -m src.main agent \
  "Find Python files in the current directory, then search for 'LocalCodingAgent' and summarize the matches." \
  --cwd ..

6.5 Shell commands

python3 -m src.main agent \
  "Run pwd and ls in the current working directory, then summarize the result." \
  --cwd .. \
  --allow-shell \
  --show-transcript

6.6 Unsafe shell mode

Use only if you intentionally want destructive-shell permission enabled.

python3 -m src.main agent \
  "Explain whether destructive shell commands are currently allowed." \
  --cwd .. \
  --allow-shell \
  --unsafe

6.7 Hook/policy tool blocking

Update the policy to block bash:

cat > ./test_cases/.claw-policy.json <<'EOF'
{
  "trusted": false,
  "denyTools": ["bash"],
  "hooks": {
    "beforePrompt": ["Respect workspace policy before acting."],
    "afterTurn": ["Persist the policy decision after each turn."]
  }
}
EOF

Then test the block:

python3 -m src.main agent \
  "Try to run bash and then explain what was blocked." \
  --cwd ./test_cases \
  --allow-shell \
  --show-transcript

Look for:

  • hook_policy_tool_block
  • tool_permission_denial
  • plugin_tool_runtime messages that now also include hook/policy guidance when present

6.8 Plan tools

python3 -m src.main agent \
  "Use update_plan to store these steps: inspect the repository, implement the change, run verification. Mark the first step in_progress and sync the plan to tasks." \
  --cwd ./test_cases \
  --allow-write \
  --show-transcript

python3 -m src.main agent "/plan" --cwd ./test_cases
python3 -m src.main agent "/tasks" --cwd ./test_cases

6.9 Remote tools

python3 -m src.main agent \
  "List the configured remote profiles, connect to staging, then report the active remote status." \
  --cwd ./test_cases \
  --show-transcript

7. Session Persistence And Resume

7.1 Create a saved session

python3 -m src.main agent \
  "Create a short TODO file in the current directory and explain what you wrote." \
  --cwd ./test_cases \
  --allow-write

At the end of the run, note:

session_id=...
session_path=...

7.2 Resume a saved session

python3 -m src.main agent-resume \
  <session-id> \
  "Continue the previous task and improve the file." \
  --allow-write \
  --show-transcript

7.3 Inspect saved sessions

ls -lt .port_sessions/agent

8. Background Sessions

8.1 Launch a background session

Use a local slash-command prompt first so you can verify the background workflow without depending on the model backend:

python3 -m src.main agent-bg "/help" --cwd ./test_cases

This prints:

  • background_id=...
  • pid=...
  • log_path=...
  • record_path=...

8.2 List background sessions

python3 -m src.main agent-ps

8.3 Read background logs

python3 -m src.main agent-logs <background-id>
python3 -m src.main agent-logs <background-id> --tail 40

8.4 Attach to the current output snapshot

python3 -m src.main agent-attach <background-id>
python3 -m src.main agent-attach <background-id> --tail 40

8.5 Kill a running background session

python3 -m src.main agent-kill <background-id>

8.6 Daemon-style wrappers

python3 -m src.main daemon start "/help" --cwd ./test_cases
python3 -m src.main daemon ps
python3 -m src.main daemon logs <background-id>
python3 -m src.main daemon attach <background-id>
python3 -m src.main daemon kill <background-id>

9. Remote Runtime CLI

9.1 Remote mode commands

python3 -m src.main remote-mode staging --cwd ./test_cases
python3 -m src.main ssh-mode staging --cwd ./test_cases
python3 -m src.main teleport-mode preview --cwd ./test_cases
python3 -m src.main direct-connect-mode direct://workspace --cwd ./test_cases
python3 -m src.main deep-link-mode preview --cwd ./test_cases

9.2 Inspect and clear remote state

python3 -m src.main remote-status --cwd ./test_cases
python3 -m src.main remote-profiles --cwd ./test_cases
python3 -m src.main remote-disconnect --cwd ./test_cases
python3 -m src.main remote-status --cwd ./test_cases

10. Structured Output / JSON Schema

Create a schema file:

cat > /tmp/claw_schema.json <<'EOF'
{
  "type": "object",
  "properties": {
    "status": { "type": "string" },
    "summary": { "type": "string" }
  },
  "required": ["status", "summary"],
  "additionalProperties": false
}
EOF

Run the agent with schema mode:

python3 -m src.main agent \
  "Inspect the current repository and respond in the requested JSON format." \
  --cwd .. \
  --response-schema-file /tmp/claw_schema.json \
  --response-schema-name claw_summary \
  --response-schema-strict

11. Budgets And Limits

11.1 Total token budget

python3 -m src.main agent \
  "Give a very long answer about the current repository." \
  --cwd .. \
  --max-total-tokens 50

11.2 Input / output token budgets

python3 -m src.main agent \
  "Read several files and explain them in detail." \
  --cwd .. \
  --max-input-tokens 80 \
  --max-output-tokens 80

11.3 Reasoning-token budget

python3 -m src.main agent \
  "Solve a multi-step task and explain the result." \
  --cwd .. \
  --max-reasoning-tokens 10

11.4 Tool-call budget

python3 -m src.main agent \
  "Read multiple files, search for symbols, and summarize the repo." \
  --cwd .. \
  --max-tool-calls 1

11.5 Delegated-task budget

python3 -m src.main agent \
  "Delegate two subtasks to inspect and summarize the repo." \
  --cwd .. \
  --max-delegated-tasks 1

11.6 Cost budget

python3 -m src.main agent \
  "Inspect the current repository and summarize it." \
  --cwd .. \
  --input-cost-per-million 0.15 \
  --output-cost-per-million 0.60 \
  --max-budget-usd 0.000001

11.7 Model-call budget

python3 -m src.main agent \
  "Continue inspecting the repository until you are done." \
  --cwd .. \
  --max-model-calls 1

11.8 Session-turn budget

python3 -m src.main agent \
  "Work through the repository across multiple turns and keep going." \
  --cwd .. \
  --max-session-turns 1

11.9 Budget overrides from local policy

cat > ./test_cases/.claw-policy.json <<'EOF'
{
  "budget": {
    "max_model_calls": 0
  }
}
EOF

python3 -m src.main agent \
  "Say hello once." \
  --cwd ./test_cases

Expected result: the run stops with a model-call budget exceeded message even though you did not pass --max-model-calls on the CLI.

12. Streaming, Continuation, And Context Reduction

12.1 Streaming assistant output

python3 -m src.main agent \
  "Produce a long explanation of the current repository architecture." \
  --cwd .. \
  --stream \
  --show-transcript

12.2 Automatic continuation after truncation

Use a small output budget so the backend is more likely to stop early:

python3 -m src.main agent \
  "Write a long, structured explanation of the current repository." \
  --cwd .. \
  --max-output-tokens 32 \
  --show-transcript

12.3 Snipping older context

python3 -m src.main agent \
  "Read claw-code-agent/src/agent_runtime.py, claw-code-agent/src/agent_session.py, claw-code-agent/src/query_engine.py, and claw-code-agent/src/plugin_runtime.py, then summarize all of them in detail." \
  --cwd .. \
  --auto-snip-threshold 120 \
  --compact-preserve-messages 0 \
  --show-transcript

12.4 Compaction boundaries

python3 -m src.main agent \
  "Read several large files from claw-code-agent/src and keep explaining the repository until the context gets compacted." \
  --cwd .. \
  --auto-compact-threshold 120 \
  --compact-preserve-messages 1 \
  --show-transcript

13. File History Replay

13.1 Create file history

python3 -m src.main agent \
  "Create notes.txt with one line, then update that line to mention file history." \
  --cwd ./test_cases \
  --allow-write

13.2 Resume and inspect replay

python3 -m src.main agent-resume \
  <session-id> \
  "Continue the previous work and tell me what files were changed before this turn." \
  --allow-write \
  --show-transcript

Look for file_history_replay messages in the transcript.

14. Nested Delegation

14.1 Basic delegated subtask

python3 -m src.main agent \
  "Delegate a subtask to inspect claw-code-agent/src/agent_runtime.py and return the summary." \
  --cwd .. \
  --show-transcript

14.2 Multiple delegated subtasks

python3 -m src.main agent \
  "Delegate one subtask to scan the repository and another to summarize it after the scan." \
  --cwd .. \
  --show-transcript

14.3 Resume a delegated child session

  1. Seed a normal saved session:
python3 -m src.main agent \
  "Inspect claw-code-agent/src/agent_tools.py and give a short summary." \
  --cwd ..
  1. Copy that session_id, then run:
python3 -m src.main agent \
  "Delegate a subtask that resumes session <session-id> and continues it." \
  --cwd .. \
  --show-transcript

14.4 Topological dependency batches

python3 -m src.main agent \
  "Delegate two subtasks: one named scan, and one named summarize that depends on scan. Use topological batching and then return the final summary." \
  --cwd .. \
  --show-transcript

Look for:

  • delegate_batch_result
  • delegate_group_result
  • batch_index=...

15. Plugin Runtime

Create a local plugin manifest:

mkdir -p ./test_cases/plugins/demo
cat > ./test_cases/plugins/demo/plugin.json <<'EOF'
{
  "name": "demo-plugin",
  "hooks": {
    "beforePrompt": "Inject plugin prompt guidance.",
    "afterTurn": "Attach plugin after-turn guidance.",
    "onResume": "Reapply plugin state on resume.",
    "beforePersist": "Persist plugin state before saving.",
    "beforeDelegate": "Add plugin guidance before delegated children run.",
    "afterDelegate": "Add plugin guidance after delegated children finish."
  },
  "toolAliases": [
    {
      "name": "plugin_read",
      "baseTool": "read_file",
      "description": "Plugin alias for reading files."
    }
  ],
  "virtualTools": [
    {
      "name": "demo_virtual",
      "description": "Return a rendered plugin response.",
      "responseTemplate": "plugin topic: {topic}"
    }
  ],
  "toolHooks": {
    "read_file": {
      "beforeTool": "Validate the path before reading.",
      "afterResult": "Summarize the file before the next action."
    }
  }
}
EOF

15.1 Plugin prompt/context discovery

python3 -m src.main agent-prompt --cwd ./test_cases
python3 -m src.main agent-context-raw --cwd ./test_cases

15.2 Plugin alias tool

echo "hello plugin" > ./test_cases/hello.txt
python3 -m src.main agent \
  "Use the plugin_read tool to read hello.txt and summarize it." \
  --cwd ./test_cases \
  --show-transcript

15.3 Plugin virtual tool

python3 -m src.main agent \
  "Use the demo_virtual tool with topic plugins and return the result." \
  --cwd ./test_cases \
  --show-transcript

15.4 Plugin before/after tool guidance

python3 -m src.main agent \
  "Read hello.txt and follow the plugin guidance around the read_file tool." \
  --cwd ./test_cases \
  --show-transcript

15.5 Plugin lifecycle with resume/persist

  1. Start a session:
python3 -m src.main agent \
  "Use the plugin system and create a saved session." \
  --cwd ./test_cases
  1. Resume it:
python3 -m src.main agent-resume \
  <session-id> \
  "Continue and mention any plugin lifecycle guidance you received." \
  --show-transcript

Look for:

  • plugin_before_persist
  • Plugin resume hooks:
  • Plugin runtime state:

16. MCP Runtime

Create a local MCP manifest:

mkdir -p ./test_cases_mcp
printf 'mcp notes\n' > ./test_cases_mcp/notes.txt
cat > ./test_cases_mcp/.claw-mcp.json <<'EOF'
{
  "servers": [
    {
      "name": "workspace",
      "resources": [
        {
          "uri": "mcp://workspace/notes",
          "name": "Notes",
          "path": "notes.txt",
          "mimeType": "text/plain"
        },
        {
          "uri": "mcp://workspace/inline",
          "name": "Inline",
          "text": "inline body"
        }
      ]
    }
  ]
}
EOF

16.1 MCP context and slash commands

python3 -m src.main agent "/mcp" --cwd ./test_cases_mcp
python3 -m src.main agent "/resources" --cwd ./test_cases_mcp
python3 -m src.main agent "/resource mcp://workspace/notes" --cwd ./test_cases_mcp
python3 -m src.main agent "/mcp (MCP)" --cwd ./test_cases_mcp
python3 -m src.main agent-context-raw --cwd ./test_cases_mcp
python3 -m src.main agent-prompt --cwd ./test_cases_mcp

16.2 MCP tools through the model loop

python3 -m src.main agent \
  "List the available MCP resources, then read mcp://workspace/notes and summarize it." \
  --cwd ./test_cases_mcp \
  --show-transcript

16.3 Read inline MCP resources

python3 -m src.main agent \
  "Read the MCP resource mcp://workspace/inline and repeat its content." \
  --cwd ./test_cases_mcp \
  --show-transcript

16.4 Real stdio MCP server transport

Create a simple stdio MCP server:

cat > ./test_cases_mcp/fake_stdio_mcp.py <<'EOF'
import json
import sys

RESOURCES = [
    {
        "uri": "mcp://remote/notes",
        "name": "Remote Notes",
        "mimeType": "text/plain",
    }
]
TOOLS = [
    {
        "name": "echo",
        "description": "Echo text",
        "inputSchema": {
            "type": "object",
            "properties": {
                "text": {"type": "string"}
            }
        }
    }
]

for raw in sys.stdin:
    raw = raw.strip()
    if not raw:
        continue
    message = json.loads(raw)
    method = message.get("method")
    if method == "initialize":
        print(json.dumps({
            "jsonrpc": "2.0",
            "id": message["id"],
            "result": {
                "protocolVersion": "2025-11-25",
                "capabilities": {"resources": {}, "tools": {}},
                "serverInfo": {"name": "fake-remote", "version": "1.0.0"}
            }
        }), flush=True)
        continue
    if method == "notifications/initialized":
        continue
    if method == "resources/list":
        print(json.dumps({
            "jsonrpc": "2.0",
            "id": message["id"],
            "result": {"resources": RESOURCES}
        }), flush=True)
        continue
    if method == "resources/read":
        uri = message.get("params", {}).get("uri")
        print(json.dumps({
            "jsonrpc": "2.0",
            "id": message["id"],
            "result": {
                "contents": [
                    {
                        "uri": uri,
                        "mimeType": "text/plain",
                        "text": "remote notes via stdio"
                    }
                ]
            }
        }), flush=True)
        continue
    if method == "tools/list":
        print(json.dumps({
            "jsonrpc": "2.0",
            "id": message["id"],
            "result": {"tools": TOOLS}
        }), flush=True)
        continue
    if method == "tools/call":
        text = message.get("params", {}).get("arguments", {}).get("text", "")
        print(json.dumps({
            "jsonrpc": "2.0",
            "id": message["id"],
            "result": {
                "content": [{"type": "text", "text": "echo:" + text}],
                "isError": False
            }
        }), flush=True)
        continue
EOF

Add the stdio server to the MCP manifest:

cat > ./test_cases_mcp/.claw-mcp.json <<'EOF'
{
  "servers": [
    {
      "name": "workspace",
      "resources": [
        {
          "uri": "mcp://workspace/notes",
          "name": "Notes",
          "path": "notes.txt",
          "mimeType": "text/plain"
        },
        {
          "uri": "mcp://workspace/inline",
          "name": "Inline",
          "text": "inline body"
        }
      ]
    }
  ],
  "mcpServers": {
    "remote": {
      "command": "python3",
      "args": ["-u", "./fake_stdio_mcp.py"]
    }
  }
}
EOF

16.5 MCP transport CLI commands

python3 -m src.main mcp-status --cwd ./test_cases_mcp
python3 -m src.main mcp-resources --cwd ./test_cases_mcp
python3 -m src.main mcp-resource mcp://remote/notes --cwd ./test_cases_mcp
python3 -m src.main mcp-tools --cwd ./test_cases_mcp
python3 -m src.main mcp-call-tool echo --arguments-json '{"text":"hello"}' --cwd ./test_cases_mcp

16.6 MCP transport slash commands

python3 -m src.main agent "/mcp" --cwd ./test_cases_mcp
python3 -m src.main agent "/mcp tools" --cwd ./test_cases_mcp
python3 -m src.main agent "/mcp tool echo" --cwd ./test_cases_mcp
python3 -m src.main agent "/resources" --cwd ./test_cases_mcp
python3 -m src.main agent "/resource mcp://remote/notes" --cwd ./test_cases_mcp

16.7 MCP transport tools through the model loop

python3 -m src.main agent \
  "List the available MCP tools, call the remote echo tool with text=hello, then summarize the result." \
  --cwd ./test_cases_mcp \
  --show-transcript

17. Task Runtime

Create a clean task workspace:

mkdir -p ./test_cases_tasks
rm -rf ./test_cases_tasks/.port_sessions

17.1 Task slash commands

python3 -m src.main agent "/tasks" --cwd ./test_cases_tasks
python3 -m src.main agent "/todo" --cwd ./test_cases_tasks
python3 -m src.main agent "/task missing-task-id" --cwd ./test_cases_tasks
python3 -m src.main agent-context-raw --cwd ./test_cases_tasks
python3 -m src.main agent-prompt --cwd ./test_cases_tasks

17.2 Create and update tasks through the model loop

python3 -m src.main agent \
  "Create a task called Review runtime tasks, then list the current tasks." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

Then inspect the stored task file:

cat ./test_cases_tasks/.port_sessions/task_runtime.json

17.3 Replace the todo list

python3 -m src.main agent \
  "Replace the current todo list with three tasks: inspect runtime, verify tests, and update docs. Mark inspect runtime as done and the others as todo." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

17.4 Read back task state

python3 -m src.main agent "/tasks" --cwd ./test_cases_tasks
python3 -m src.main agent "/task-next" --cwd ./test_cases_tasks
python3 -m src.main agent \
  "List the current tasks and show me the id of each one." \
  --cwd ./test_cases_tasks \
  --show-transcript

17.5 Plan runtime and task sync

python3 -m src.main agent \
  "Use update_plan to create three steps: inspect runtime, verify tests, update docs. Mark inspect runtime completed and sync to tasks." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

python3 -m src.main agent "/plan" --cwd ./test_cases_tasks
python3 -m src.main agent "/tasks" --cwd ./test_cases_tasks
cat ./test_cases_tasks/.port_sessions/plan_runtime.json

17.6 Dependency-aware task execution

Create a blocked task graph through the model loop:

python3 -m src.main agent \
  "Use todo_write to create two tasks: scan with status pending, and patch with status blocked and blocked_by scan. Then show the next actionable tasks." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

Then advance the task state:

python3 -m src.main agent \
  "Mark task scan as completed, then show the next actionable tasks and start task patch with active_form 'Patching files'." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

Inspect the resulting task state:

python3 -m src.main agent "/task-next" --cwd ./test_cases_tasks
python3 -m src.main agent "/tasks" --cwd ./test_cases_tasks
cat ./test_cases_tasks/.port_sessions/task_runtime.json

17.7 Task execution tools directly

python3 -m src.main agent \
  "Use todo_write to create task scan and task patch where patch is blocked_by scan. Then use task_next, task_complete for scan, and task_start for patch." \
  --cwd ./test_cases_tasks \
  --allow-write \
  --show-transcript

18. Query Engine And Workspace Commands

18.1 Workspace inventory

python3 -m src.main summary
python3 -m src.main manifest
python3 -m src.main subsystems --limit 20
python3 -m src.main commands --limit 20
python3 -m src.main tools --limit 20

18.2 Query routing and bootstrap reports

python3 -m src.main route "inspect the runtime and tools" --limit 10
python3 -m src.main bootstrap "inspect the runtime and tools" --limit 10
python3 -m src.main turn-loop "inspect the runtime and tools" --limit 5 --max-turns 3

18.3 Session flushing for the mirrored workspace

python3 -m src.main flush-transcript "store a temporary transcript"
python3 -m src.main load-session <session-id>

19. Parity Tracking Workflow

Use this every time a new feature lands:

python3 -m unittest discover -s tests -v

Then update:

  • PARITY_CHECKLIST.md
  • TESTING_GUIDE.md

Rule for future work:

  • every new implemented feature should add a checked item in PARITY_CHECKLIST.md
  • every user-testable feature should add at least one concrete command example in TESTING_GUIDE.md

20. Config Runtime

Inspect local config state:

python3 -m src.main config-status --cwd ./test_cases
python3 -m src.main config-effective --cwd ./test_cases

Read a value or a source directly:

python3 -m src.main config-get review.mode --cwd ./test_cases
python3 -m src.main config-source project --cwd ./test_cases

Write a value into the local settings layer:

python3 -m src.main config-set review.mode '"strict"' --cwd ./test_cases
python3 -m src.main config-set review.enabled true --cwd ./test_cases

Test the slash commands:

python3 -m src.main agent "/config" --cwd ./test_cases
python3 -m src.main agent "/config effective" --cwd ./test_cases
python3 -m src.main agent "/config get review.mode" --cwd ./test_cases
python3 -m src.main agent "/settings source local" --cwd ./test_cases

Test the config tools through the real agent loop:

python3 -m src.main agent \
  "List the current config keys, set review.mode to strict in local config, then read it back." \
  --cwd ./test_cases \
  --allow-write

21. Account Runtime

Inspect local account state:

python3 -m src.main account-status --cwd ./test_cases
python3 -m src.main account-profiles --cwd ./test_cases

Activate and clear a local account session:

python3 -m src.main account-login local --cwd ./test_cases
python3 -m src.main account-logout --cwd ./test_cases

Test the slash commands:

python3 -m src.main agent "/account" --cwd ./test_cases
python3 -m src.main agent "/account profiles" --cwd ./test_cases
python3 -m src.main agent "/login local" --cwd ./test_cases
python3 -m src.main agent "/logout" --cwd ./test_cases

Test the account tools through the real agent loop:

python3 -m src.main agent \
  "List the configured account profiles, activate the local profile, then report the active account session." \
  --cwd ./test_cases

22. Extended Tool Slice

Test the new tool-surface additions through the real agent loop:

python3 -m src.main agent \
  "Use tool_search to find file-related tools and summarize the best ones for reading and editing files." \
  --cwd ./test_cases
python3 -m src.main agent \
  "Use web_fetch on file://$(pwd)/README.md and summarize the first section." \
  --cwd .
python3 -m src.main agent \
  "Call the sleep tool for 0.1 seconds, then tell me it completed." \
  --cwd ./test_cases

Run the direct unit tests for this slice:

python3 -m unittest tests.test_extended_tools -v