AI systems get easier when abstraction is high and data is abundant. Natural language and math have pretraining support and clean symbolic structure. They sit in the blessed region.
Physical AI, robotics, neuro-data, world models, cognitive security, and research metadata sit closer to cursed regions: lower data, weaker ground truth, harder compression, unstable objectives.
Thesis
Useful model evaluation starts with measurable tasks.
When data and ground truth are scarce, defining the task and its checks can matter more than model selection.
The map
High-abstraction, high-data tasks can borrow structure from pretraining.
Lower-data or lower-abstraction domains need explicit scaffolds: object identity, typed relations, state boundaries, hard validators, trace capture, and human review loops.
Why blessed-region methods fail elsewhere
A final reward can obscure which decision caused success or failure.
Capability jumps can make a method look causal when the model crossed a threshold where the task became compressible.
Agent systems make the issue sharper. Post-training depends on whether the environment produces localized feedback.
Tool-call failures, state-boundary violations, runtime errors, and verifier failures should become correction signals, not terminal noise spread across whole rollouts.
What survives in cursed regions
- Exact artifacts and preserved provenance.
- Stable object identity and typed relations.
- State boundaries, hard validators, and calibrated uncertainty.
- Trace capture, review loops, and operator-facing evidence.
- Learned residuals only after the measurable environment exists.
How this shows up in my work
T-UEBA used constrained graph ML, active learning, synthetic-data triggers, and calibration because tactical behavior does not stay stationary.
DL-Patcher and the separate VULCAN/AIxCC effort used patch curation, code-model harnesses, and verification.
DXF-to-CAD research explores geometric objects, relation graphs, and learned repair under deterministic checks.
Practical rule
Build the environment and measurement loop first. Spend compute only where the feedback signal is real.
Research map · Agent post-training as environment design · Selected work