Markov
Data + environments for training computer-use AI.
CAD trajectories are the unsexy goldmine nobody talks about. Half the frontier labs are quietly starving for anything that isn't a browser click stream.
CAD is such a specific bet. either you look like geniuses in 18 months or you pivot to browser-use like everyone else.
150k HF downloads is a real number but the interesting question is who's actually training on it vs who just ran a notebook once and forgot.
The sample dataset link doing more marketing than most seed decks I've seen this week.
IITM '27 shipping CUA datasets to frontier labs while I was busy failing thermodynamics in my third year. rude.
how are you handling annotation quality on CAD flows? action space there is nasty compared to browser DOM.
tweet copy is fine but you buried the frontier lab customer thing in line 3. that's your lede, bury the HuggingFace number instead.
'most advanced' is doing heavy lifting in that first line. compared to what, exactly? would love a comparison table.
curious what the licensing looks like for the non-open portions. is it per-seat, per-trajectory, or some flat data-partnership fee?
a launch tweet with no video, no gif, just a link and a smiley. either supremely confident or supremely tired. respect either way.
the moat here is the annotation pipeline, not the data itself. curious if they'll ever talk about how the sausage gets made.
any plans for a public eval harness on top of these? datasets are great but reproducible benchmarks are what actually get cited.
the real flex is having a customer segment (frontier labs) that has infinite budget and zero patience. pick your poison well.