docs: publish architecture and experiment portal

This commit is contained in:
wuyang
2026-07-26 21:03:23 +08:00
parent 4eb915283d
commit 688ddec1fe
17 changed files with 3582 additions and 56 deletions
+223
View File
@@ -0,0 +1,223 @@
# Agent loop implementation
## Purpose
The K1412 loop turns a model completion API into an accountable coding Agent.
Its primary rule is simple: a textual claim is not proof that work happened.
Files, commands, reports, and tests must be backed by successful tool events
from the current run.
The current event metadata identifies the implementation as:
- strategy: `agent-loop-v3`;
- scheduler: `safe-parallel-v1`;
- context policy: `recent-visible-v1`.
These identifiers are part of the experiment contract, not marketing version
numbers. Change them when behavior changes in a way that can affect evaluation.
## Run state machine
```mermaid
stateDiagram-v2
[*] --> BuildContext
BuildContext --> RequestModel
RequestModel --> ValidateCalls: tool calls
RequestModel --> CheckCompletion: no tool calls
ValidateCalls --> Schedule
Schedule --> Execute
Execute --> RecordEvidence
RecordEvidence --> RequestModel
CheckCompletion --> RequestModel: evidence missing and budget remains
CheckCompletion --> Completed: evidence sufficient
CheckCompletion --> Unverified: evidence missing and retry budget exhausted
Completed --> [*]
Unverified --> [*]
```
Every top-level run receives a unique `run_id`. All events include the user,
chat, monotonic sequence, timestamp, event type, and a sanitized payload.
## Context policy
`ContextPolicy.prepare`:
1. accepts only system, developer, user, and assistant messages;
2. strips rendered `<details type="tool_calls">` blocks from old assistant
messages so UI markup is not fed back as model context;
3. prepends the K1412 system contract;
4. appends up to eight durable user memories as context, explicitly not as
higher-priority instructions;
5. walks visible messages newest-first until the model's character budget is
reached;
6. records how many old messages were dropped.
The policy is intentionally simple and inspectable. It does not yet summarize
old conversation branches, retrieve semantic workspace context, or estimate
provider-specific tokenization. Those are future experiment candidates.
## Model contract
`ModelSpec` is the server-side source of truth for:
- public and provider model IDs;
- provider selection;
- thinking capability and display label;
- reasoning effort when the provider supports it;
- maximum output tokens;
- maximum loop iterations;
- context character budget.
Luna, Terra, and Sol are separate local models with boolean thinking support;
they are not three effort settings for one model. DeepSeek V4 Pro uses the
DeepSeek provider with thinking enabled and `reasoning_effort=max`.
Runtime preserves provider `reasoning_content` across tool turns when required
for protocol correctness, but does not publish hidden reasoning as user-facing
content.
## Tool catalog
Workspace tools:
- status, list, read, search, write, and patch files;
- execute foreground commands;
- inspect Git status and diff;
- start, poll, and cancel background processes.
Runtime state tools:
- update a per-chat plan;
- remember, recall, and forget durable user memory.
Delegation tool:
- start one bounded child Agent with a role, task, and explicit write policy.
Tool schemas use `additionalProperties: false` so malformed model arguments
fail early. Arguments are normalized before policy checks. Content and patch
bodies are omitted from public run-event payloads.
## Intent-based tool selection
The root loop does not always send every tool. Lightweight request classifiers
detect whether the task concerns artifacts, execution, source files, reports,
comparisons, processes, Git, memory, or delegation. The selected catalog keeps
the model's tool decision smaller while retaining the core workspace tools.
This is heuristic routing, not authorization. Gateway remains the enforcement
boundary.
## Scheduler
Each model response can request at most eight tool calls. Calls are parsed in
their original order and divided into consecutive groups:
- tools marked `parallel_safe` execute concurrently;
- mutations and other non-parallel tools execute serially;
- read-only child Agents may run in parallel;
- a child with write access is serialized;
- child Agents cannot delegate again.
Gateway also serializes workspace mutations per user. This second lock is
important because multiple Runtime requests or browser actions can address the
same workspace.
The scheduler preserves result order even when a parallel group completes out
of order.
## Defensive normalization and retry guards
The loop repairs a small set of common model formatting mistakes, then applies
policy:
- reject bare interactive commands such as `python3`, `bash`, or `node`;
- reject attempts to run documentation/data files as executable source;
- reject syntactically invalid complete Python files before writing;
- decode repeated escaped source-layout newlines produced by weak models;
- skip identical successful writes;
- block an identical failed write until content changes;
- block an unchanged failed command until a mutation or materially different
diagnostic occurs;
- deduplicate repeated commands in one batch;
- cap generation, iteration, and tool-batch sizes.
At the execution layer, Bash runs with `pipefail` and returns the true exit
code. Foreground tools default to a 900-second ceiling; background processes
have explicit start/poll/cancel lifecycle.
## Evidence-gated completion
The loop derives required evidence from the user request.
| Request shape | Minimum evidence |
| --- | --- |
| Concrete artifact | A successful file mutation and a later successful verification |
| Execute/test/analyze | At least one successful execution or inspection |
| Source-code request | A separate executable source file |
| Report request | A separate report written after execution |
| Benchmark/comparison report | Exact numeric measurements copied from successful execution output |
Verification includes a later file read, Git diff/status, command execution, or
background-process poll. A failed command remains unresolved until a later
execution succeeds.
If a model tries to finish too early, Runtime emits `completion.rejected`,
returns a focused recovery instruction to the model, and continues. After
repeated rejected checkpoints or iteration exhaustion, Runtime returns an
explicit incomplete result instead of laundering an unverified claim into a
success.
## Delegation
A root Agent can delegate a bounded independent task. The child receives:
- a role and exact task;
- a reduced iteration budget of at most eight;
- no delegation tool;
- read-only tools unless write access was explicitly requested.
Child events share the root run stream with an increased `depth`. The current
implementation is recursive execution inside one Runtime process, not a
distributed queue or a persistent autonomous worker.
## Memory and plans
Durable memories are scoped by user ID. Recall is currently recency-first with
optional case-insensitive substring filtering. Plans are scoped by user and
chat and allow at most one `in_progress` item.
Memory is useful but intentionally conservative: the system does not
automatically extract facts from every conversation and does not yet perform
embedding retrieval, confidence scoring, expiry, or conflict resolution.
## Run event model
Important event types include:
- `run.created`, `context.built`, `run.completed`, `run.failed`,
`run.cancelled`;
- `model.requested`, `model.responded`;
- `tool.started`, `tool.completed`, `tool.batch_limited`;
- `completion.rejected`, `completion.unverified`;
- `agent.spawned`, `agent.completed`.
The event stream powers UI detail blocks today and is the basis for future
replay, evaluation, cost analysis, and A/B assignment. Events store public
arguments and concise summaries, not credentials, full write bodies, or hidden
reasoning.
## Known limitations
- Context compaction drops old messages instead of summarizing them.
- Memory retrieval is lexical rather than semantic.
- Tool intent classifiers are regular-expression based.
- The scheduler parallelizes only consecutive safe calls and has no resource
cost model.
- Child Agents are not durable across Runtime restarts.
- Experiment assignment and aggregate dashboards are not yet implemented.
- Provider token accounting is retained when returned but is not yet converted
into a unified cost model.
These are deliberate experiment targets, documented in
[experiments.md](experiments.md).