Files
zk-data-agent/benchmarks/README.md
T
copilot-swe-agent[bot] 231b977b92 Add 10 standard evaluation benchmark suites with CLI runner and README
Implements HumanEval, MBPP, SWE-Bench, Aider, LiveCodeBench (coding),
MATH, GSM8K, AIME (math), and IFEval, BFCL (instruction following).

Each suite includes built-in problem subsets (108 total) and supports
loading full datasets from JSONL files. Includes comprehensive README
with all commands.

Agent-Logs-Url: https://github.com/HarnessLab/claw-code-agent/sessions/6890e3d0-3058-4b1f-b7e5-27171c079c62

Co-authored-by: abdoelsayed2016 <27821589+abdoelsayed2016@users.noreply.github.com>
2026-04-05 19:58:00 +00:00

334 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Benchmarks — claw-code-agent
This directory contains two benchmark systems:
1. **Local task benchmarks** (`benchmarks/run.py`) — custom tasks that test agent capabilities directly
2. **Standard evaluation suites** (`benchmarks/run_suite.py`) — implementations of well-known AI evaluation benchmarks
---
## Quick Start
```bash
# From the repository root:
# Set up your model endpoint
export OPENAI_API_KEY="your-key"
export OPENAI_MODEL="gpt-4" # or your model name
export OPENAI_BASE_URL="http://localhost:8000/v1" # if using local vLLM/ollama
# Run a quick smoke test (5 problems from HumanEval)
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
# Run all benchmarks
python3 -m benchmarks.run_suite --all -o results.json
```
---
## Standard Evaluation Suites
### Available Benchmarks
| Suite | Category | # Problems (built-in) | Description |
|-------|----------|----------------------|-------------|
| **HumanEval** | Coding | 20 | Code generation from Python docstrings |
| **MBPP** | Coding | 15 | Basic Python programming problems |
| **SWE-Bench** | Coding | 5 | Resolve real-world GitHub issues |
| **Aider** | Coding | 6 | Code editing and refactoring tasks |
| **LiveCodeBench** | Coding | 5 | Competitive programming problems |
| **MATH** | Math | 15 | Competition mathematics problems |
| **GSM8K** | Math | 15 | Grade school math word problems |
| **AIME** | Math | 10 | Challenging competition math (integers 0999) |
| **IFEval** | Instruction Following | 10 | Verifiable instruction-following evaluation |
| **BFCL** | Instruction Following | 7 | Function/tool calling evaluation |
### Commands
#### List All Suites
```bash
python3 -m benchmarks.run_suite --list
```
#### Run a Specific Suite
```bash
# Coding benchmarks
python3 -m benchmarks.run_suite --suite humaneval
python3 -m benchmarks.run_suite --suite mbpp
python3 -m benchmarks.run_suite --suite swe-bench
python3 -m benchmarks.run_suite --suite aider
python3 -m benchmarks.run_suite --suite livecodebench
# Math benchmarks
python3 -m benchmarks.run_suite --suite math
python3 -m benchmarks.run_suite --suite gsm8k
python3 -m benchmarks.run_suite --suite aime
# Instruction following benchmarks
python3 -m benchmarks.run_suite --suite ifeval
python3 -m benchmarks.run_suite --suite bfcl
```
#### Run by Category
```bash
# All coding benchmarks (~51 problems)
python3 -m benchmarks.run_suite --category coding
# All math benchmarks (~40 problems)
python3 -m benchmarks.run_suite --category math
# All instruction following benchmarks (~17 problems)
python3 -m benchmarks.run_suite --category instruction-following
```
#### Run ALL Suites
```bash
python3 -m benchmarks.run_suite --all
```
#### Limit Problems (Quick Testing)
```bash
# Run only first 3 problems from each suite
python3 -m benchmarks.run_suite --all --limit 3
# Quick HumanEval smoke test
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
```
#### Save Results to JSON
```bash
python3 -m benchmarks.run_suite --all -o results.json
python3 -m benchmarks.run_suite --suite humaneval -o humaneval_results.json
```
#### Verbose Output
```bash
python3 -m benchmarks.run_suite --suite humaneval -v
```
#### Custom Timeout
```bash
# 10 minutes per problem
python3 -m benchmarks.run_suite --suite swe-bench --timeout 600
```
#### Use Full Datasets (JSONL)
```bash
# Download or place your datasets in benchmarks/data/
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./benchmarks/data/
# Or specify a custom directory
python3 -m benchmarks.run_suite --suite humaneval --data-dir /path/to/datasets/
```
---
## Using Full Datasets
Each suite looks for a JSONL file in the data directory. If the file isn't found, it falls back to the built-in subset.
| Suite | Expected File | Format |
|-------|--------------|--------|
| HumanEval | `humaneval.jsonl` | `{"task_id", "prompt", "canonical_solution", "test", "entry_point"}` |
| MBPP | `mbpp.jsonl` | `{"task_id", "text", "code", "test_list"}` |
| SWE-Bench | `swe_bench.jsonl` | `{"instance_id", "problem_statement", "setup_code", "test_cmd"}` |
| Aider | `aider.jsonl` | `{"id", "instruction", "setup_code", "test_code"}` |
| LiveCodeBench | `livecodebench.jsonl` | `{"id", "title", "description", "test_cases", "function_name"}` |
| MATH | `math.jsonl` | `{"id", "problem", "answer", "subject", "level"}` |
| GSM8K | `gsm8k.jsonl` | `{"id", "question", "answer"}` |
| AIME | `aime.jsonl` | `{"id", "problem", "answer"}` |
| IFEval | `ifeval.jsonl` | `{"id", "instruction", "checks"}` |
| BFCL | `bfcl.jsonl` | `{"id", "instruction", "expected_function", "setup_code", "test_code"}` |
### Downloading Full Datasets
```bash
# HumanEval (from OpenAI)
wget -O benchmarks/data/humaneval.jsonl \
https://raw.githubusercontent.com/openai/human-eval/master/data/HumanEval.jsonl.gz
gunzip benchmarks/data/humaneval.jsonl.gz
# GSM8K (from HuggingFace — requires datasets library)
pip install datasets
python3 -c "
from datasets import load_dataset
import json
ds = load_dataset('gsm8k', 'main', split='test')
with open('benchmarks/data/gsm8k.jsonl', 'w') as f:
for i, item in enumerate(ds):
f.write(json.dumps({'id': f'gsm8k-{i}', 'question': item['question'], 'answer': item['answer'].split('####')[-1].strip()}) + '\n')
"
# MATH (from HuggingFace)
python3 -c "
from datasets import load_dataset
import json
ds = load_dataset('hendrycks/competition_math', split='test')
with open('benchmarks/data/math.jsonl', 'w') as f:
for i, item in enumerate(ds):
f.write(json.dumps({'id': f'math-{i}', 'problem': item['problem'], 'answer': item['solution'], 'subject': item['type'], 'level': item['level']}) + '\n')
"
```
---
## Local Task Benchmarks
The original local benchmark system tests the agent on custom tasks.
```bash
# Run all local tasks
python3 -m benchmarks.run
# Run a single task
python3 -m benchmarks.run --task file-create-basic
# Run a category
python3 -m benchmarks.run --category bugfix
# Run by difficulty
python3 -m benchmarks.run --difficulty easy
# List available tasks
python3 -m benchmarks.run --list
# Verbose + save results
python3 -m benchmarks.run -v -o local_results.json
```
---
## All Commands Reference
```bash
# ─── STANDARD EVALUATION SUITES ────────────────────────────────────────
# List suites
python3 -m benchmarks.run_suite --list
# Individual suites
python3 -m benchmarks.run_suite --suite humaneval # Coding: docstring → code
python3 -m benchmarks.run_suite --suite mbpp # Coding: basic Python
python3 -m benchmarks.run_suite --suite swe-bench # Coding: GitHub issues
python3 -m benchmarks.run_suite --suite aider # Coding: code editing
python3 -m benchmarks.run_suite --suite livecodebench # Coding: competitive programming
python3 -m benchmarks.run_suite --suite math # Math: competition math
python3 -m benchmarks.run_suite --suite gsm8k # Math: grade school
python3 -m benchmarks.run_suite --suite aime # Math: AIME competition
python3 -m benchmarks.run_suite --suite ifeval # Instruction: format following
python3 -m benchmarks.run_suite --suite bfcl # Instruction: function calling
# Category runs
python3 -m benchmarks.run_suite --category coding # All coding (~51 problems)
python3 -m benchmarks.run_suite --category math # All math (~40 problems)
python3 -m benchmarks.run_suite --category instruction-following # All IF (~17 problems)
# Full run
python3 -m benchmarks.run_suite --all # All suites (~108 problems)
python3 -m benchmarks.run_suite --all --limit 3 # Quick: 3 per suite
python3 -m benchmarks.run_suite --all -v -o results.json # Full + verbose + save
# Options
python3 -m benchmarks.run_suite --suite humaneval --limit 5 # Limit problems
python3 -m benchmarks.run_suite --suite humaneval --timeout 600 # Custom timeout (seconds)
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./data # Custom data directory
python3 -m benchmarks.run_suite --suite humaneval -v # Verbose mode
python3 -m benchmarks.run_suite --suite humaneval -o out.json # Save to JSON
# ─── LOCAL TASK BENCHMARKS ─────────────────────────────────────────────
python3 -m benchmarks.run --list # List local tasks
python3 -m benchmarks.run # Run all local tasks
python3 -m benchmarks.run --task file-create-basic # Single task
python3 -m benchmarks.run --category bugfix # Category filter
python3 -m benchmarks.run --difficulty easy # Difficulty filter
python3 -m benchmarks.run -v -o local_results.json # Verbose + save
```
---
## Environment Variables
| Variable | Description | Example |
|----------|-------------|---------|
| `OPENAI_API_KEY` | API key for the model provider | `sk-...` |
| `OPENAI_MODEL` | Model name to use | `gpt-4`, `qwen2.5-coder-32b` |
| `OPENAI_BASE_URL` | API base URL (for local models) | `http://localhost:8000/v1` |
---
## Output Format
Results are saved as JSON with this structure:
```json
{
"benchmark_run": "claw-code-agent-evaluation",
"timestamp": "2025-01-15T10:30:00",
"suites": [
{
"suite_name": "HumanEval",
"total": 20,
"passed": 15,
"failed": 5,
"score_pct": 75.0,
"duration_sec": 450.5,
"model": "gpt-4",
"results": [
{
"problem_id": "HumanEval/0",
"passed": true,
"duration_sec": 12.3,
"error": ""
}
]
}
],
"summary": {
"total_suites": 10,
"total_problems": 108,
"total_passed": 85,
"overall_score_pct": 78.7
}
}
```
---
## Architecture
```
benchmarks/
├── __init__.py
├── run.py # Local task benchmark runner
├── run_suite.py # Standard evaluation suite runner (CLI)
├── README.md # This file
├── data/ # Dataset files (JSONL) — not committed
│ └── .gitkeep
├── tasks/
│ ├── __init__.py
│ └── definitions.py # Local task definitions
└── suites/
├── __init__.py
├── base.py # Base class for all suites
├── humaneval.py # HumanEval benchmark
├── mbpp.py # MBPP benchmark
├── swe_bench.py # SWE-Bench benchmark
├── aider.py # Aider benchmark
├── livecodebench.py # LiveCodeBench benchmark
├── math_bench.py # MATH benchmark
├── gsm8k.py # GSM8K benchmark
├── aime.py # AIME benchmark
├── ifeval.py # IFEval benchmark
└── bfcl.py # BFCL benchmark
```