231b977b92
Implements HumanEval, MBPP, SWE-Bench, Aider, LiveCodeBench (coding), MATH, GSM8K, AIME (math), and IFEval, BFCL (instruction following). Each suite includes built-in problem subsets (108 total) and supports loading full datasets from JSONL files. Includes comprehensive README with all commands. Agent-Logs-Url: https://github.com/HarnessLab/claw-code-agent/sessions/6890e3d0-3058-4b1f-b7e5-27171c079c62 Co-authored-by: abdoelsayed2016 <27821589+abdoelsayed2016@users.noreply.github.com>
334 lines
11 KiB
Markdown
334 lines
11 KiB
Markdown
# Benchmarks — claw-code-agent
|
||
|
||
This directory contains two benchmark systems:
|
||
|
||
1. **Local task benchmarks** (`benchmarks/run.py`) — custom tasks that test agent capabilities directly
|
||
2. **Standard evaluation suites** (`benchmarks/run_suite.py`) — implementations of well-known AI evaluation benchmarks
|
||
|
||
---
|
||
|
||
## Quick Start
|
||
|
||
```bash
|
||
# From the repository root:
|
||
|
||
# Set up your model endpoint
|
||
export OPENAI_API_KEY="your-key"
|
||
export OPENAI_MODEL="gpt-4" # or your model name
|
||
export OPENAI_BASE_URL="http://localhost:8000/v1" # if using local vLLM/ollama
|
||
|
||
# Run a quick smoke test (5 problems from HumanEval)
|
||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
|
||
|
||
# Run all benchmarks
|
||
python3 -m benchmarks.run_suite --all -o results.json
|
||
```
|
||
|
||
---
|
||
|
||
## Standard Evaluation Suites
|
||
|
||
### Available Benchmarks
|
||
|
||
| Suite | Category | # Problems (built-in) | Description |
|
||
|-------|----------|----------------------|-------------|
|
||
| **HumanEval** | Coding | 20 | Code generation from Python docstrings |
|
||
| **MBPP** | Coding | 15 | Basic Python programming problems |
|
||
| **SWE-Bench** | Coding | 5 | Resolve real-world GitHub issues |
|
||
| **Aider** | Coding | 6 | Code editing and refactoring tasks |
|
||
| **LiveCodeBench** | Coding | 5 | Competitive programming problems |
|
||
| **MATH** | Math | 15 | Competition mathematics problems |
|
||
| **GSM8K** | Math | 15 | Grade school math word problems |
|
||
| **AIME** | Math | 10 | Challenging competition math (integers 0–999) |
|
||
| **IFEval** | Instruction Following | 10 | Verifiable instruction-following evaluation |
|
||
| **BFCL** | Instruction Following | 7 | Function/tool calling evaluation |
|
||
|
||
### Commands
|
||
|
||
#### List All Suites
|
||
|
||
```bash
|
||
python3 -m benchmarks.run_suite --list
|
||
```
|
||
|
||
#### Run a Specific Suite
|
||
|
||
```bash
|
||
# Coding benchmarks
|
||
python3 -m benchmarks.run_suite --suite humaneval
|
||
python3 -m benchmarks.run_suite --suite mbpp
|
||
python3 -m benchmarks.run_suite --suite swe-bench
|
||
python3 -m benchmarks.run_suite --suite aider
|
||
python3 -m benchmarks.run_suite --suite livecodebench
|
||
|
||
# Math benchmarks
|
||
python3 -m benchmarks.run_suite --suite math
|
||
python3 -m benchmarks.run_suite --suite gsm8k
|
||
python3 -m benchmarks.run_suite --suite aime
|
||
|
||
# Instruction following benchmarks
|
||
python3 -m benchmarks.run_suite --suite ifeval
|
||
python3 -m benchmarks.run_suite --suite bfcl
|
||
```
|
||
|
||
#### Run by Category
|
||
|
||
```bash
|
||
# All coding benchmarks (~51 problems)
|
||
python3 -m benchmarks.run_suite --category coding
|
||
|
||
# All math benchmarks (~40 problems)
|
||
python3 -m benchmarks.run_suite --category math
|
||
|
||
# All instruction following benchmarks (~17 problems)
|
||
python3 -m benchmarks.run_suite --category instruction-following
|
||
```
|
||
|
||
#### Run ALL Suites
|
||
|
||
```bash
|
||
python3 -m benchmarks.run_suite --all
|
||
```
|
||
|
||
#### Limit Problems (Quick Testing)
|
||
|
||
```bash
|
||
# Run only first 3 problems from each suite
|
||
python3 -m benchmarks.run_suite --all --limit 3
|
||
|
||
# Quick HumanEval smoke test
|
||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
|
||
```
|
||
|
||
#### Save Results to JSON
|
||
|
||
```bash
|
||
python3 -m benchmarks.run_suite --all -o results.json
|
||
python3 -m benchmarks.run_suite --suite humaneval -o humaneval_results.json
|
||
```
|
||
|
||
#### Verbose Output
|
||
|
||
```bash
|
||
python3 -m benchmarks.run_suite --suite humaneval -v
|
||
```
|
||
|
||
#### Custom Timeout
|
||
|
||
```bash
|
||
# 10 minutes per problem
|
||
python3 -m benchmarks.run_suite --suite swe-bench --timeout 600
|
||
```
|
||
|
||
#### Use Full Datasets (JSONL)
|
||
|
||
```bash
|
||
# Download or place your datasets in benchmarks/data/
|
||
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./benchmarks/data/
|
||
|
||
# Or specify a custom directory
|
||
python3 -m benchmarks.run_suite --suite humaneval --data-dir /path/to/datasets/
|
||
```
|
||
|
||
---
|
||
|
||
## Using Full Datasets
|
||
|
||
Each suite looks for a JSONL file in the data directory. If the file isn't found, it falls back to the built-in subset.
|
||
|
||
| Suite | Expected File | Format |
|
||
|-------|--------------|--------|
|
||
| HumanEval | `humaneval.jsonl` | `{"task_id", "prompt", "canonical_solution", "test", "entry_point"}` |
|
||
| MBPP | `mbpp.jsonl` | `{"task_id", "text", "code", "test_list"}` |
|
||
| SWE-Bench | `swe_bench.jsonl` | `{"instance_id", "problem_statement", "setup_code", "test_cmd"}` |
|
||
| Aider | `aider.jsonl` | `{"id", "instruction", "setup_code", "test_code"}` |
|
||
| LiveCodeBench | `livecodebench.jsonl` | `{"id", "title", "description", "test_cases", "function_name"}` |
|
||
| MATH | `math.jsonl` | `{"id", "problem", "answer", "subject", "level"}` |
|
||
| GSM8K | `gsm8k.jsonl` | `{"id", "question", "answer"}` |
|
||
| AIME | `aime.jsonl` | `{"id", "problem", "answer"}` |
|
||
| IFEval | `ifeval.jsonl` | `{"id", "instruction", "checks"}` |
|
||
| BFCL | `bfcl.jsonl` | `{"id", "instruction", "expected_function", "setup_code", "test_code"}` |
|
||
|
||
### Downloading Full Datasets
|
||
|
||
```bash
|
||
# HumanEval (from OpenAI)
|
||
wget -O benchmarks/data/humaneval.jsonl \
|
||
https://raw.githubusercontent.com/openai/human-eval/master/data/HumanEval.jsonl.gz
|
||
gunzip benchmarks/data/humaneval.jsonl.gz
|
||
|
||
# GSM8K (from HuggingFace — requires datasets library)
|
||
pip install datasets
|
||
python3 -c "
|
||
from datasets import load_dataset
|
||
import json
|
||
ds = load_dataset('gsm8k', 'main', split='test')
|
||
with open('benchmarks/data/gsm8k.jsonl', 'w') as f:
|
||
for i, item in enumerate(ds):
|
||
f.write(json.dumps({'id': f'gsm8k-{i}', 'question': item['question'], 'answer': item['answer'].split('####')[-1].strip()}) + '\n')
|
||
"
|
||
|
||
# MATH (from HuggingFace)
|
||
python3 -c "
|
||
from datasets import load_dataset
|
||
import json
|
||
ds = load_dataset('hendrycks/competition_math', split='test')
|
||
with open('benchmarks/data/math.jsonl', 'w') as f:
|
||
for i, item in enumerate(ds):
|
||
f.write(json.dumps({'id': f'math-{i}', 'problem': item['problem'], 'answer': item['solution'], 'subject': item['type'], 'level': item['level']}) + '\n')
|
||
"
|
||
```
|
||
|
||
---
|
||
|
||
## Local Task Benchmarks
|
||
|
||
The original local benchmark system tests the agent on custom tasks.
|
||
|
||
```bash
|
||
# Run all local tasks
|
||
python3 -m benchmarks.run
|
||
|
||
# Run a single task
|
||
python3 -m benchmarks.run --task file-create-basic
|
||
|
||
# Run a category
|
||
python3 -m benchmarks.run --category bugfix
|
||
|
||
# Run by difficulty
|
||
python3 -m benchmarks.run --difficulty easy
|
||
|
||
# List available tasks
|
||
python3 -m benchmarks.run --list
|
||
|
||
# Verbose + save results
|
||
python3 -m benchmarks.run -v -o local_results.json
|
||
```
|
||
|
||
---
|
||
|
||
## All Commands Reference
|
||
|
||
```bash
|
||
# ─── STANDARD EVALUATION SUITES ────────────────────────────────────────
|
||
|
||
# List suites
|
||
python3 -m benchmarks.run_suite --list
|
||
|
||
# Individual suites
|
||
python3 -m benchmarks.run_suite --suite humaneval # Coding: docstring → code
|
||
python3 -m benchmarks.run_suite --suite mbpp # Coding: basic Python
|
||
python3 -m benchmarks.run_suite --suite swe-bench # Coding: GitHub issues
|
||
python3 -m benchmarks.run_suite --suite aider # Coding: code editing
|
||
python3 -m benchmarks.run_suite --suite livecodebench # Coding: competitive programming
|
||
python3 -m benchmarks.run_suite --suite math # Math: competition math
|
||
python3 -m benchmarks.run_suite --suite gsm8k # Math: grade school
|
||
python3 -m benchmarks.run_suite --suite aime # Math: AIME competition
|
||
python3 -m benchmarks.run_suite --suite ifeval # Instruction: format following
|
||
python3 -m benchmarks.run_suite --suite bfcl # Instruction: function calling
|
||
|
||
# Category runs
|
||
python3 -m benchmarks.run_suite --category coding # All coding (~51 problems)
|
||
python3 -m benchmarks.run_suite --category math # All math (~40 problems)
|
||
python3 -m benchmarks.run_suite --category instruction-following # All IF (~17 problems)
|
||
|
||
# Full run
|
||
python3 -m benchmarks.run_suite --all # All suites (~108 problems)
|
||
python3 -m benchmarks.run_suite --all --limit 3 # Quick: 3 per suite
|
||
python3 -m benchmarks.run_suite --all -v -o results.json # Full + verbose + save
|
||
|
||
# Options
|
||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 # Limit problems
|
||
python3 -m benchmarks.run_suite --suite humaneval --timeout 600 # Custom timeout (seconds)
|
||
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./data # Custom data directory
|
||
python3 -m benchmarks.run_suite --suite humaneval -v # Verbose mode
|
||
python3 -m benchmarks.run_suite --suite humaneval -o out.json # Save to JSON
|
||
|
||
# ─── LOCAL TASK BENCHMARKS ─────────────────────────────────────────────
|
||
|
||
python3 -m benchmarks.run --list # List local tasks
|
||
python3 -m benchmarks.run # Run all local tasks
|
||
python3 -m benchmarks.run --task file-create-basic # Single task
|
||
python3 -m benchmarks.run --category bugfix # Category filter
|
||
python3 -m benchmarks.run --difficulty easy # Difficulty filter
|
||
python3 -m benchmarks.run -v -o local_results.json # Verbose + save
|
||
```
|
||
|
||
---
|
||
|
||
## Environment Variables
|
||
|
||
| Variable | Description | Example |
|
||
|----------|-------------|---------|
|
||
| `OPENAI_API_KEY` | API key for the model provider | `sk-...` |
|
||
| `OPENAI_MODEL` | Model name to use | `gpt-4`, `qwen2.5-coder-32b` |
|
||
| `OPENAI_BASE_URL` | API base URL (for local models) | `http://localhost:8000/v1` |
|
||
|
||
---
|
||
|
||
## Output Format
|
||
|
||
Results are saved as JSON with this structure:
|
||
|
||
```json
|
||
{
|
||
"benchmark_run": "claw-code-agent-evaluation",
|
||
"timestamp": "2025-01-15T10:30:00",
|
||
"suites": [
|
||
{
|
||
"suite_name": "HumanEval",
|
||
"total": 20,
|
||
"passed": 15,
|
||
"failed": 5,
|
||
"score_pct": 75.0,
|
||
"duration_sec": 450.5,
|
||
"model": "gpt-4",
|
||
"results": [
|
||
{
|
||
"problem_id": "HumanEval/0",
|
||
"passed": true,
|
||
"duration_sec": 12.3,
|
||
"error": ""
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"summary": {
|
||
"total_suites": 10,
|
||
"total_problems": 108,
|
||
"total_passed": 85,
|
||
"overall_score_pct": 78.7
|
||
}
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## Architecture
|
||
|
||
```
|
||
benchmarks/
|
||
├── __init__.py
|
||
├── run.py # Local task benchmark runner
|
||
├── run_suite.py # Standard evaluation suite runner (CLI)
|
||
├── README.md # This file
|
||
├── data/ # Dataset files (JSONL) — not committed
|
||
│ └── .gitkeep
|
||
├── tasks/
|
||
│ ├── __init__.py
|
||
│ └── definitions.py # Local task definitions
|
||
└── suites/
|
||
├── __init__.py
|
||
├── base.py # Base class for all suites
|
||
├── humaneval.py # HumanEval benchmark
|
||
├── mbpp.py # MBPP benchmark
|
||
├── swe_bench.py # SWE-Bench benchmark
|
||
├── aider.py # Aider benchmark
|
||
├── livecodebench.py # LiveCodeBench benchmark
|
||
├── math_bench.py # MATH benchmark
|
||
├── gsm8k.py # GSM8K benchmark
|
||
├── aime.py # AIME benchmark
|
||
├── ifeval.py # IFEval benchmark
|
||
└── bfcl.py # BFCL benchmark
|
||
```
|