Abdelrahman Abdallah
|
aacf0a212a
|
Created src/prompt_constants.py (550+ lines) porting all constants from npm src/constants/:
┌──────────────────┬───────────────────────────┬───────────────────────────────────────────────┐
│ Category │ npm Source │ Items │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Product metadata │ product.ts │ URLs, base URLs │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ System prefixes │ system.ts │ 3 prompt prefixes │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Cyber risk │ cyberRiskInstruction.ts │ Safety instruction │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ API limits │ apiLimits.ts │ 10 image/PDF/media limits │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Tool limits │ toolLimits.ts │ 6 result size constants │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Spinner verbs │ spinnerVerbs.ts │ 187 whimsical gerunds │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Completion verbs │ turnCompletionVerbs.ts │ 8 past-tense verbs │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Figures/symbols │ figures.ts │ 25 Unicode UI symbols │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ XML tags │ xml.ts │ 30+ tag constants │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Messages │ messages.ts │ NO_CONTENT_MESSAGE │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Date utilities │ common.ts │ 4 functions │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Section caching │ systemPromptSections.ts │ Memoized/volatile sections │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Output styles │ outputStyles.ts │ 3 built-in configs │
├──────────────────┼───────────────────────────┼───────────────────────────────────────────────┤
│ Prompt helpers │ prompts.ts │ Knowledge cutoff, language, scratchpad, hooks │
└──────────────────┴───────────────────────────┴───────────────────────────────────────────────┘
91 new tests in tests/test_prompt_constants.py. All 17 SQL todos done.
|
2026-04-08 00:02:53 +02:00 |
|
Abdelrahman Abdallah
|
90489e7bfc
|
Add 7 new benchmark suites for Gemma 4 comparison
New suites:
- MMLU-Pro: Professional-level 10-choice QA (14 subjects)
- GPQA-Diamond: Graduate-level science QA
- BigBench Extra Hard: Challenging reasoning tasks
- MMMLU: Multilingual MMLU across 10 languages
- HLE: Humanity's Last Exam (extremely hard)
- Tau2: Tool-augmented reasoning (retail/airline/finance)
- Codeforces: Competitive programming with ELO scoring
Updates:
- AIME now supports aime_2026.jsonl for 2026 problems
- Registry expanded to 17 suites in 5 categories
- download_datasets.py supports HF downloads for new suites
- README.md updated with full documentation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
2026-04-07 04:55:31 +02:00 |
|
Abdelrahman Abdallah
|
a5629295ac
|
Implemented the next parity slice.
New runtime/code:
- src/ask_user_runtime.py
- src/team_runtime.py
New real tools in src/agent_tools.py:
- ask_user_question
- team_create
- team_delete
- team_list
- team_get
- send_message
- team_messages
- notebook_edit
|
2026-04-07 02:51:30 +02:00 |
|
Abdelrahman Abdallah
|
a54c90b18f
|
update the codebase and clean up it
|
2026-04-06 03:42:44 +02:00 |
|
copilot-swe-agent[bot]
|
589a74f4a8
|
Fix review comments: remove empty string concat, fix variable scoping in BFCL evaluate
Agent-Logs-Url: https://github.com/HarnessLab/claw-code-agent/sessions/6890e3d0-3058-4b1f-b7e5-27171c079c62
Co-authored-by: abdoelsayed2016 <27821589+abdoelsayed2016@users.noreply.github.com>
|
2026-04-05 19:59:16 +00:00 |
|
copilot-swe-agent[bot]
|
231b977b92
|
Add 10 standard evaluation benchmark suites with CLI runner and README
Implements HumanEval, MBPP, SWE-Bench, Aider, LiveCodeBench (coding),
MATH, GSM8K, AIME (math), and IFEval, BFCL (instruction following).
Each suite includes built-in problem subsets (108 total) and supports
loading full datasets from JSONL files. Includes comprehensive README
with all commands.
Agent-Logs-Url: https://github.com/HarnessLab/claw-code-agent/sessions/6890e3d0-3058-4b1f-b7e5-27171c079c62
Co-authored-by: abdoelsayed2016 <27821589+abdoelsayed2016@users.noreply.github.com>
|
2026-04-05 19:58:00 +00:00 |
|
Abdelrahman Abdallah
|
783145fe6a
|
add mcp and online search
|
2026-04-05 02:35:49 +02:00 |
|