Self-Scaffolding LLMs
Self-scaffolding LLMs learn task-specific scaffolds or harnesses alongside solution rollouts. Ornith-1.0 frames the scaffold as a learnable object co-evolving with the policy for agentic coding tasks, with safeguards against reward hacking [src-182].
Related
- Agentic Coding
- Reward Hacking
- AI Evaluation
Recommended next
Keep reading from this thread
From 477 indexed pages and articles.
- Wiki concept Reward Hacking A failure mode where a model or agent optimizes the scoring signal while missing the real objective. Related by scaffolding
- Wiki concept Agentic Coding Software work performed through tool-using models that plan, edit, test, inspect errors, and iterate over codebases. Related by scaffolding
- Insight Agentic Web Readiness A practical checklist for making websites understandable, navigable, and useful for AI agents and answer engines Readers have engaged with this next