Reward Hacking
Reward hacking is a failure mode where a model or agent optimizes the scoring signal while missing the real objective.
Key facts
- Ornith-1.0 is relevant to reward-hacking scrutiny because self-generated scaffolds could improve coding performance or simply learn benchmark-specific strategies unless evaluated carefully [src-182].
Related
Recommended next
Keep reading from this thread
From 477 indexed pages and articles.
- Wiki concept Self-Scaffolding LLMs Learn task-specific scaffolds or harnesses alongside solution rollouts. Ornith-1.0 frames the scaffold as a learnable object co-evolving with the policy for agentic Related by hacking
- Wiki concept AI Evaluation Covers the methods used to measure model or agent performance, reliability, robustness, and alignment with intended outcomes. Related by hacking
- Insight AI Beyond POCs How enterprise AI moves beyond proofs of concept through ownership, governance, measurement, adoption, and production operating models Readers have engaged with this next