Agent Harness Training
A training framework that improves agent harnesses while keeping the LLM frozen, updating prompts, tools, and context management in a training loop. End-to-end determinism enables credit assignment to harness changes. Trained on a 35B open model, the harness transferred to six other held-out models on SWE-bench and Terminal-Bench 2.0 and outperformed the official Terminus 2 harness.

Python, Agent Systems, Harness Engineering, LLM Evaluation



