Add 10 standard evaluation benchmark suites with CLI runner and README
Implements HumanEval, MBPP, SWE-Bench, Aider, LiveCodeBench (coding), MATH, GSM8K, AIME (math), and IFEval, BFCL (instruction following). Each suite includes built-in problem subsets (108 total) and supports loading full datasets from JSONL files. Includes comprehensive README with all commands. Agent-Logs-Url: https://github.com/HarnessLab/claw-code-agent/sessions/6890e3d0-3058-4b1f-b7e5-27171c079c62 Co-authored-by: abdoelsayed2016 <27821589+abdoelsayed2016@users.noreply.github.com>
This commit is contained in:
committed by
GitHub
parent
3e32154618
commit
231b977b92
@@ -0,0 +1,333 @@
|
||||
# Benchmarks — claw-code-agent
|
||||
|
||||
This directory contains two benchmark systems:
|
||||
|
||||
1. **Local task benchmarks** (`benchmarks/run.py`) — custom tasks that test agent capabilities directly
|
||||
2. **Standard evaluation suites** (`benchmarks/run_suite.py`) — implementations of well-known AI evaluation benchmarks
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# From the repository root:
|
||||
|
||||
# Set up your model endpoint
|
||||
export OPENAI_API_KEY="your-key"
|
||||
export OPENAI_MODEL="gpt-4" # or your model name
|
||||
export OPENAI_BASE_URL="http://localhost:8000/v1" # if using local vLLM/ollama
|
||||
|
||||
# Run a quick smoke test (5 problems from HumanEval)
|
||||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
|
||||
|
||||
# Run all benchmarks
|
||||
python3 -m benchmarks.run_suite --all -o results.json
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Standard Evaluation Suites
|
||||
|
||||
### Available Benchmarks
|
||||
|
||||
| Suite | Category | # Problems (built-in) | Description |
|
||||
|-------|----------|----------------------|-------------|
|
||||
| **HumanEval** | Coding | 20 | Code generation from Python docstrings |
|
||||
| **MBPP** | Coding | 15 | Basic Python programming problems |
|
||||
| **SWE-Bench** | Coding | 5 | Resolve real-world GitHub issues |
|
||||
| **Aider** | Coding | 6 | Code editing and refactoring tasks |
|
||||
| **LiveCodeBench** | Coding | 5 | Competitive programming problems |
|
||||
| **MATH** | Math | 15 | Competition mathematics problems |
|
||||
| **GSM8K** | Math | 15 | Grade school math word problems |
|
||||
| **AIME** | Math | 10 | Challenging competition math (integers 0–999) |
|
||||
| **IFEval** | Instruction Following | 10 | Verifiable instruction-following evaluation |
|
||||
| **BFCL** | Instruction Following | 7 | Function/tool calling evaluation |
|
||||
|
||||
### Commands
|
||||
|
||||
#### List All Suites
|
||||
|
||||
```bash
|
||||
python3 -m benchmarks.run_suite --list
|
||||
```
|
||||
|
||||
#### Run a Specific Suite
|
||||
|
||||
```bash
|
||||
# Coding benchmarks
|
||||
python3 -m benchmarks.run_suite --suite humaneval
|
||||
python3 -m benchmarks.run_suite --suite mbpp
|
||||
python3 -m benchmarks.run_suite --suite swe-bench
|
||||
python3 -m benchmarks.run_suite --suite aider
|
||||
python3 -m benchmarks.run_suite --suite livecodebench
|
||||
|
||||
# Math benchmarks
|
||||
python3 -m benchmarks.run_suite --suite math
|
||||
python3 -m benchmarks.run_suite --suite gsm8k
|
||||
python3 -m benchmarks.run_suite --suite aime
|
||||
|
||||
# Instruction following benchmarks
|
||||
python3 -m benchmarks.run_suite --suite ifeval
|
||||
python3 -m benchmarks.run_suite --suite bfcl
|
||||
```
|
||||
|
||||
#### Run by Category
|
||||
|
||||
```bash
|
||||
# All coding benchmarks (~51 problems)
|
||||
python3 -m benchmarks.run_suite --category coding
|
||||
|
||||
# All math benchmarks (~40 problems)
|
||||
python3 -m benchmarks.run_suite --category math
|
||||
|
||||
# All instruction following benchmarks (~17 problems)
|
||||
python3 -m benchmarks.run_suite --category instruction-following
|
||||
```
|
||||
|
||||
#### Run ALL Suites
|
||||
|
||||
```bash
|
||||
python3 -m benchmarks.run_suite --all
|
||||
```
|
||||
|
||||
#### Limit Problems (Quick Testing)
|
||||
|
||||
```bash
|
||||
# Run only first 3 problems from each suite
|
||||
python3 -m benchmarks.run_suite --all --limit 3
|
||||
|
||||
# Quick HumanEval smoke test
|
||||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 -v
|
||||
```
|
||||
|
||||
#### Save Results to JSON
|
||||
|
||||
```bash
|
||||
python3 -m benchmarks.run_suite --all -o results.json
|
||||
python3 -m benchmarks.run_suite --suite humaneval -o humaneval_results.json
|
||||
```
|
||||
|
||||
#### Verbose Output
|
||||
|
||||
```bash
|
||||
python3 -m benchmarks.run_suite --suite humaneval -v
|
||||
```
|
||||
|
||||
#### Custom Timeout
|
||||
|
||||
```bash
|
||||
# 10 minutes per problem
|
||||
python3 -m benchmarks.run_suite --suite swe-bench --timeout 600
|
||||
```
|
||||
|
||||
#### Use Full Datasets (JSONL)
|
||||
|
||||
```bash
|
||||
# Download or place your datasets in benchmarks/data/
|
||||
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./benchmarks/data/
|
||||
|
||||
# Or specify a custom directory
|
||||
python3 -m benchmarks.run_suite --suite humaneval --data-dir /path/to/datasets/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Using Full Datasets
|
||||
|
||||
Each suite looks for a JSONL file in the data directory. If the file isn't found, it falls back to the built-in subset.
|
||||
|
||||
| Suite | Expected File | Format |
|
||||
|-------|--------------|--------|
|
||||
| HumanEval | `humaneval.jsonl` | `{"task_id", "prompt", "canonical_solution", "test", "entry_point"}` |
|
||||
| MBPP | `mbpp.jsonl` | `{"task_id", "text", "code", "test_list"}` |
|
||||
| SWE-Bench | `swe_bench.jsonl` | `{"instance_id", "problem_statement", "setup_code", "test_cmd"}` |
|
||||
| Aider | `aider.jsonl` | `{"id", "instruction", "setup_code", "test_code"}` |
|
||||
| LiveCodeBench | `livecodebench.jsonl` | `{"id", "title", "description", "test_cases", "function_name"}` |
|
||||
| MATH | `math.jsonl` | `{"id", "problem", "answer", "subject", "level"}` |
|
||||
| GSM8K | `gsm8k.jsonl` | `{"id", "question", "answer"}` |
|
||||
| AIME | `aime.jsonl` | `{"id", "problem", "answer"}` |
|
||||
| IFEval | `ifeval.jsonl` | `{"id", "instruction", "checks"}` |
|
||||
| BFCL | `bfcl.jsonl` | `{"id", "instruction", "expected_function", "setup_code", "test_code"}` |
|
||||
|
||||
### Downloading Full Datasets
|
||||
|
||||
```bash
|
||||
# HumanEval (from OpenAI)
|
||||
wget -O benchmarks/data/humaneval.jsonl \
|
||||
https://raw.githubusercontent.com/openai/human-eval/master/data/HumanEval.jsonl.gz
|
||||
gunzip benchmarks/data/humaneval.jsonl.gz
|
||||
|
||||
# GSM8K (from HuggingFace — requires datasets library)
|
||||
pip install datasets
|
||||
python3 -c "
|
||||
from datasets import load_dataset
|
||||
import json
|
||||
ds = load_dataset('gsm8k', 'main', split='test')
|
||||
with open('benchmarks/data/gsm8k.jsonl', 'w') as f:
|
||||
for i, item in enumerate(ds):
|
||||
f.write(json.dumps({'id': f'gsm8k-{i}', 'question': item['question'], 'answer': item['answer'].split('####')[-1].strip()}) + '\n')
|
||||
"
|
||||
|
||||
# MATH (from HuggingFace)
|
||||
python3 -c "
|
||||
from datasets import load_dataset
|
||||
import json
|
||||
ds = load_dataset('hendrycks/competition_math', split='test')
|
||||
with open('benchmarks/data/math.jsonl', 'w') as f:
|
||||
for i, item in enumerate(ds):
|
||||
f.write(json.dumps({'id': f'math-{i}', 'problem': item['problem'], 'answer': item['solution'], 'subject': item['type'], 'level': item['level']}) + '\n')
|
||||
"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Local Task Benchmarks
|
||||
|
||||
The original local benchmark system tests the agent on custom tasks.
|
||||
|
||||
```bash
|
||||
# Run all local tasks
|
||||
python3 -m benchmarks.run
|
||||
|
||||
# Run a single task
|
||||
python3 -m benchmarks.run --task file-create-basic
|
||||
|
||||
# Run a category
|
||||
python3 -m benchmarks.run --category bugfix
|
||||
|
||||
# Run by difficulty
|
||||
python3 -m benchmarks.run --difficulty easy
|
||||
|
||||
# List available tasks
|
||||
python3 -m benchmarks.run --list
|
||||
|
||||
# Verbose + save results
|
||||
python3 -m benchmarks.run -v -o local_results.json
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## All Commands Reference
|
||||
|
||||
```bash
|
||||
# ─── STANDARD EVALUATION SUITES ────────────────────────────────────────
|
||||
|
||||
# List suites
|
||||
python3 -m benchmarks.run_suite --list
|
||||
|
||||
# Individual suites
|
||||
python3 -m benchmarks.run_suite --suite humaneval # Coding: docstring → code
|
||||
python3 -m benchmarks.run_suite --suite mbpp # Coding: basic Python
|
||||
python3 -m benchmarks.run_suite --suite swe-bench # Coding: GitHub issues
|
||||
python3 -m benchmarks.run_suite --suite aider # Coding: code editing
|
||||
python3 -m benchmarks.run_suite --suite livecodebench # Coding: competitive programming
|
||||
python3 -m benchmarks.run_suite --suite math # Math: competition math
|
||||
python3 -m benchmarks.run_suite --suite gsm8k # Math: grade school
|
||||
python3 -m benchmarks.run_suite --suite aime # Math: AIME competition
|
||||
python3 -m benchmarks.run_suite --suite ifeval # Instruction: format following
|
||||
python3 -m benchmarks.run_suite --suite bfcl # Instruction: function calling
|
||||
|
||||
# Category runs
|
||||
python3 -m benchmarks.run_suite --category coding # All coding (~51 problems)
|
||||
python3 -m benchmarks.run_suite --category math # All math (~40 problems)
|
||||
python3 -m benchmarks.run_suite --category instruction-following # All IF (~17 problems)
|
||||
|
||||
# Full run
|
||||
python3 -m benchmarks.run_suite --all # All suites (~108 problems)
|
||||
python3 -m benchmarks.run_suite --all --limit 3 # Quick: 3 per suite
|
||||
python3 -m benchmarks.run_suite --all -v -o results.json # Full + verbose + save
|
||||
|
||||
# Options
|
||||
python3 -m benchmarks.run_suite --suite humaneval --limit 5 # Limit problems
|
||||
python3 -m benchmarks.run_suite --suite humaneval --timeout 600 # Custom timeout (seconds)
|
||||
python3 -m benchmarks.run_suite --suite humaneval --data-dir ./data # Custom data directory
|
||||
python3 -m benchmarks.run_suite --suite humaneval -v # Verbose mode
|
||||
python3 -m benchmarks.run_suite --suite humaneval -o out.json # Save to JSON
|
||||
|
||||
# ─── LOCAL TASK BENCHMARKS ─────────────────────────────────────────────
|
||||
|
||||
python3 -m benchmarks.run --list # List local tasks
|
||||
python3 -m benchmarks.run # Run all local tasks
|
||||
python3 -m benchmarks.run --task file-create-basic # Single task
|
||||
python3 -m benchmarks.run --category bugfix # Category filter
|
||||
python3 -m benchmarks.run --difficulty easy # Difficulty filter
|
||||
python3 -m benchmarks.run -v -o local_results.json # Verbose + save
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Environment Variables
|
||||
|
||||
| Variable | Description | Example |
|
||||
|----------|-------------|---------|
|
||||
| `OPENAI_API_KEY` | API key for the model provider | `sk-...` |
|
||||
| `OPENAI_MODEL` | Model name to use | `gpt-4`, `qwen2.5-coder-32b` |
|
||||
| `OPENAI_BASE_URL` | API base URL (for local models) | `http://localhost:8000/v1` |
|
||||
|
||||
---
|
||||
|
||||
## Output Format
|
||||
|
||||
Results are saved as JSON with this structure:
|
||||
|
||||
```json
|
||||
{
|
||||
"benchmark_run": "claw-code-agent-evaluation",
|
||||
"timestamp": "2025-01-15T10:30:00",
|
||||
"suites": [
|
||||
{
|
||||
"suite_name": "HumanEval",
|
||||
"total": 20,
|
||||
"passed": 15,
|
||||
"failed": 5,
|
||||
"score_pct": 75.0,
|
||||
"duration_sec": 450.5,
|
||||
"model": "gpt-4",
|
||||
"results": [
|
||||
{
|
||||
"problem_id": "HumanEval/0",
|
||||
"passed": true,
|
||||
"duration_sec": 12.3,
|
||||
"error": ""
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"total_suites": 10,
|
||||
"total_problems": 108,
|
||||
"total_passed": 85,
|
||||
"overall_score_pct": 78.7
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
benchmarks/
|
||||
├── __init__.py
|
||||
├── run.py # Local task benchmark runner
|
||||
├── run_suite.py # Standard evaluation suite runner (CLI)
|
||||
├── README.md # This file
|
||||
├── data/ # Dataset files (JSONL) — not committed
|
||||
│ └── .gitkeep
|
||||
├── tasks/
|
||||
│ ├── __init__.py
|
||||
│ └── definitions.py # Local task definitions
|
||||
└── suites/
|
||||
├── __init__.py
|
||||
├── base.py # Base class for all suites
|
||||
├── humaneval.py # HumanEval benchmark
|
||||
├── mbpp.py # MBPP benchmark
|
||||
├── swe_bench.py # SWE-Bench benchmark
|
||||
├── aider.py # Aider benchmark
|
||||
├── livecodebench.py # LiveCodeBench benchmark
|
||||
├── math_bench.py # MATH benchmark
|
||||
├── gsm8k.py # GSM8K benchmark
|
||||
├── aime.py # AIME benchmark
|
||||
├── ifeval.py # IFEval benchmark
|
||||
└── bfcl.py # BFCL benchmark
|
||||
```
|
||||
Reference in New Issue
Block a user