Agent training and evaluation depend on the tasks, tools, observations, and feedback available to the model.
Feedback
A final success score can hide the cause of failure. Tool errors, rejected patches, and state changes help locate the decision that needs correction.
Scoring
A useful evaluation tests whether an agent completes the task and whether it can exploit the scoring rule. Inspecting traces and testing negative cases helps distinguish the two.
Related work
- DL-Patcher and VULCAN/AIxCC: patch curation, code-model harnesses, and verification in separate research efforts.
- T-UEBA: calibration and analyst feedback for changing tactical behavior.
- Robot planning: reproducible comparisons and checks on the complete execution path.