Prime Intellect
Own Your Intelligence
About
self-modifiable harness state is a fun phrase until the harness decides your test suite was the problem.
ok wait, programmatic tool calling plus context as a variable is actually the part I want to see benchmarked, not the demo where it fixes a typo in a README.
we shipped something eerily similar as an internal harness in 2019, we just called it a Jenkins job with anxiety.
before I say interesting, what does week 4 retention look like on the coding tasks? self-improving is a strong claim to underwrite.
the launch page pacing is clean but the tweet copy reads like a NeurIPS abstract someone squeezed into 280 chars.
hot take: shipping this as a uv tool install one-liner does more marketing work than the whole video did.
every coding agent harness launch feels like we're all speedrunning the same plateau, the ceiling on this category is lower than the pitch decks suggest.
reminds me of what a portco of mine tried to build on top of Claude last year, except this one might actually ship past the demo.
curious what the token economics look like when the harness is rewriting itself mid-task, feels like a fun way to torch a compute budget.
how many people actually built this? because if it's under ten I'm going to need a moment.
self-modifiable harness plus long-running autonomy is going to be a fun conversation with any enterprise security team, hope there's an audit log story somewhere.
who's hiring RL infra folks off the back of this? got a candidate wrapping up at a robotics lab who eats harnesses for breakfast.