Benchmarks — claw-code-agent
This directory contains two benchmark systems:
- Local task benchmarks (
benchmarks/run.py) — custom tasks that test agent capabilities directly - Standard evaluation suites (
benchmarks/run_suite.py) — implementations of well-known AI evaluation benchmarks
Quick Start
# From the repository root:
# Set up your model endpoint
export OPENAI_API_KEY="your-key"
export OPENAI_MODEL="gpt-4" # or your model name
export OPENAI_BASE_URL="http://localhost:8000/v1" # if using local vLLM/ollama
# Run a quick smoke test (5 problems from HumanEval)
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
# Run all benchmarks
python3 -m benchmarks.run_suite --all -o results.json
Standard Evaluation Suites
Available Benchmarks
| Suite | Category | # Problems (built-in) | Description |
|---|---|---|---|
| HumanEval | Coding | 20 | Code generation from Python docstrings |
| MBPP | Coding | 15 | Basic Python programming problems |
| SWE-Bench | Coding | 5 | Resolve real-world GitHub issues |
| Aider | Coding | 6 | Code editing and refactoring tasks |
| LiveCodeBench | Coding | 5 | Competitive programming problems |
| MATH | Math | 15 | Competition mathematics problems |
| GSM8K | Math | 15 | Grade school math word problems |
| AIME | Math | 10 | Challenging competition math (integers 0–999) |
| IFEval | Instruction Following | 10 | Verifiable instruction-following evaluation |
| BFCL | Instruction Following | 7 | Function/tool calling evaluation |
Commands
List All Suites
python3 -m benchmarks.run_suite --list
Run a Specific Suite
# Coding benchmarks
python3 -m benchmarks.run_suite --suite humaneval
python3 -m benchmarks.run_suite --suite mbpp
python3 -m benchmarks.run_suite --suite swe-bench
python3 -m benchmarks.run_suite --suite aider
python3 -m benchmarks.run_suite --suite livecodebench
# Math benchmarks
python3 -m benchmarks.run_suite --suite math
python3 -m benchmarks.run_suite --suite gsm8k
python3 -m benchmarks.run_suite --suite aime
# Instruction following benchmarks
python3 -m benchmarks.run_suite --suite ifeval
python3 -m benchmarks.run_suite --suite bfcl
Run by Category
# All coding benchmarks (~51 problems)
python3 -m benchmarks.run_suite --category coding
# All math benchmarks (~40 problems)
python3 -m benchmarks.run_suite --category math
# All instruction following benchmarks (~17 problems)
python3 -m benchmarks.run_suite --category instruction-following
Run ALL Suites
python3 -m benchmarks.run_suite --all
Limit Problems (Quick Testing)
# Run only first 3 problems from each suite
python3 -m benchmarks.run_suite --all --limit 3
# Quick HumanEval smoke test
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
Save Results to JSON
python3 -m benchmarks.run_suite --all -o results.json
python3 -m benchmarks.run_suite --suite humaneval -o humaneval_results.json
Verbose Output
python3 -m benchmarks.run_suite --suite humaneval -v
Custom Timeout
# 10 minutes per problem
python3 -m benchmarks.run_suite --suite swe-bench --timeout 600
Use Full Datasets (JSONL)
# Download or place your datasets in benchmarks/data/
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./benchmarks/data/
# Or specify a custom directory
python3 -m benchmarks.run_suite --suite humaneval --data-dir /path/to/datasets/
Using Full Datasets
Each suite looks for a JSONL file in the data directory. If the file isn't found, it falls back to the built-in subset.
| Suite | Expected File | Format |
|---|---|---|
| HumanEval | humaneval.jsonl |
{"task_id", "prompt", "canonical_solution", "test", "entry_point"} |
| MBPP | mbpp.jsonl |
{"task_id", "text", "code", "test_list"} |
| SWE-Bench | swe_bench.jsonl |
{"instance_id", "problem_statement", "setup_code", "test_cmd"} |
| Aider | aider.jsonl |
{"id", "instruction", "setup_code", "test_code"} |
| LiveCodeBench | livecodebench.jsonl |
{"id", "title", "description", "test_cases", "function_name"} |
| MATH | math.jsonl |
{"id", "problem", "answer", "subject", "level"} |
| GSM8K | gsm8k.jsonl |
{"id", "question", "answer"} |
| AIME | aime.jsonl |
{"id", "problem", "answer"} |
| IFEval | ifeval.jsonl |
{"id", "instruction", "checks"} |
| BFCL | bfcl.jsonl |
{"id", "instruction", "expected_function", "setup_code", "test_code"} |
Downloading Full Datasets
# HumanEval (from OpenAI)
wget -O benchmarks/data/humaneval.jsonl \
https://raw.githubusercontent.com/openai/human-eval/master/data/HumanEval.jsonl.gz
gunzip benchmarks/data/humaneval.jsonl.gz
# GSM8K (from HuggingFace — requires datasets library)
pip install datasets
python3 -c "
from datasets import load_dataset
import json
ds = load_dataset('gsm8k', 'main', split='test')
with open('benchmarks/data/gsm8k.jsonl', 'w') as f:
for i, item in enumerate(ds):
f.write(json.dumps({'id': f'gsm8k-{i}', 'question': item['question'], 'answer': item['answer'].split('####')[-1].strip()}) + '\n')
"
# MATH (from HuggingFace)
python3 -c "
from datasets import load_dataset
import json
ds = load_dataset('hendrycks/competition_math', split='test')
with open('benchmarks/data/math.jsonl', 'w') as f:
for i, item in enumerate(ds):
f.write(json.dumps({'id': f'math-{i}', 'problem': item['problem'], 'answer': item['solution'], 'subject': item['type'], 'level': item['level']}) + '\n')
"
Local Task Benchmarks
The original local benchmark system tests the agent on custom tasks.
# Run all local tasks
python3 -m benchmarks.run
# Run a single task
python3 -m benchmarks.run --task file-create-basic
# Run a category
python3 -m benchmarks.run --category bugfix
# Run by difficulty
python3 -m benchmarks.run --difficulty easy
# List available tasks
python3 -m benchmarks.run --list
# Verbose + save results
python3 -m benchmarks.run -v -o local_results.json
All Commands Reference
# ─── STANDARD EVALUATION SUITES ────────────────────────────────────────
# List suites
python3 -m benchmarks.run_suite --list
# Individual suites
python3 -m benchmarks.run_suite --suite humaneval # Coding: docstring → code
python3 -m benchmarks.run_suite --suite mbpp # Coding: basic Python
python3 -m benchmarks.run_suite --suite swe-bench # Coding: GitHub issues
python3 -m benchmarks.run_suite --suite aider # Coding: code editing
python3 -m benchmarks.run_suite --suite livecodebench # Coding: competitive programming
python3 -m benchmarks.run_suite --suite math # Math: competition math
python3 -m benchmarks.run_suite --suite gsm8k # Math: grade school
python3 -m benchmarks.run_suite --suite aime # Math: AIME competition
python3 -m benchmarks.run_suite --suite ifeval # Instruction: format following
python3 -m benchmarks.run_suite --suite bfcl # Instruction: function calling
# Category runs
python3 -m benchmarks.run_suite --category coding # All coding (~51 problems)
python3 -m benchmarks.run_suite --category math # All math (~40 problems)
python3 -m benchmarks.run_suite --category instruction-following # All IF (~17 problems)
# Full run
python3 -m benchmarks.run_suite --all # All suites (~108 problems)
python3 -m benchmarks.run_suite --all --limit 3 # Quick: 3 per suite
python3 -m benchmarks.run_suite --all -v -o results.json # Full + verbose + save
# Options
python3 -m benchmarks.run_suite --suite humaneval --limit 5 # Limit problems
python3 -m benchmarks.run_suite --suite humaneval --timeout 600 # Custom timeout (seconds)
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./data # Custom data directory
python3 -m benchmarks.run_suite --suite humaneval -v # Verbose mode
python3 -m benchmarks.run_suite --suite humaneval -o out.json # Save to JSON
# ─── LOCAL TASK BENCHMARKS ─────────────────────────────────────────────
python3 -m benchmarks.run --list # List local tasks
python3 -m benchmarks.run # Run all local tasks
python3 -m benchmarks.run --task file-create-basic # Single task
python3 -m benchmarks.run --category bugfix # Category filter
python3 -m benchmarks.run --difficulty easy # Difficulty filter
python3 -m benchmarks.run -v -o local_results.json # Verbose + save
Environment Variables
| Variable | Description | Example |
|---|---|---|
OPENAI_API_KEY |
API key for the model provider | sk-... |
OPENAI_MODEL |
Model name to use | gpt-4, qwen2.5-coder-32b |
OPENAI_BASE_URL |
API base URL (for local models) | http://localhost:8000/v1 |
Output Format
Results are saved as JSON with this structure:
{
"benchmark_run": "claw-code-agent-evaluation",
"timestamp": "2025-01-15T10:30:00",
"suites": [
{
"suite_name": "HumanEval",
"total": 20,
"passed": 15,
"failed": 5,
"score_pct": 75.0,
"duration_sec": 450.5,
"model": "gpt-4",
"results": [
{
"problem_id": "HumanEval/0",
"passed": true,
"duration_sec": 12.3,
"error": ""
}
]
}
],
"summary": {
"total_suites": 10,
"total_problems": 108,
"total_passed": 85,
"overall_score_pct": 78.7
}
}
Architecture
benchmarks/
├── __init__.py
├── run.py # Local task benchmark runner
├── run_suite.py # Standard evaluation suite runner (CLI)
├── README.md # This file
├── data/ # Dataset files (JSONL) — not committed
│ └── .gitkeep
├── tasks/
│ ├── __init__.py
│ └── definitions.py # Local task definitions
└── suites/
├── __init__.py
├── base.py # Base class for all suites
├── humaneval.py # HumanEval benchmark
├── mbpp.py # MBPP benchmark
├── swe_bench.py # SWE-Bench benchmark
├── aider.py # Aider benchmark
├── livecodebench.py # LiveCodeBench benchmark
├── math_bench.py # MATH benchmark
├── gsm8k.py # GSM8K benchmark
├── aime.py # AIME benchmark
├── ifeval.py # IFEval benchmark
└── bfcl.py # BFCL benchmark