Back to Timeline

Event Summary

Epoch AI and METR published MirrorCode results on long-horizon coding tasks that require models to reimplement programs without source-code or web access. One Claude Opus 4.7 run completed a 16,000-line target in 14 hours at a reported $251 inference cost.

Impact Assessment

  • Capability Leap +2 · Medium-term

    MirrorCode provides evidence that current models can complete some long-horizon program-reimplementation tasks under controlled conditions.

    Affected Groups: software engineers, AI-agent researchers, developer-tool builders

Consensus & Sources

Significance L1
Category Capability Breakthrough
Consensus Emerging Consensus
Impact Index 4/10