Back to directory
Enterprise infrastructure · Data platforms & warehouses

Hanji

Turn complex documents into AI-ready data

building Hanji (YC X25). built @youlearnai for millions of students.
746 followers
TLVC Rating
Hook
Editing / Creativity
Copy
Sentiment of launch
Distribution strategy
Community Rating
No ratings yet
Your rating
Sign in to rate this launch.

About

Hanji is an API for turning messy source documents (scans, faxes, handwritten forms, degraded government paperwork, dense tables) into structured data that downstream AI systems can actually use. The company is aimed at healthcare, insurance, and government teams , groups that tend to either maintain brittle in-house parsing pipelines or accept fixed tradeoffs between accuracy, cost, and latency from off-the-shelf providers. Hanji's pitch is that those three dials can be tuned per customer, with the vendor doing the tuning against the customer's actual document mix. The launch matters now because document extraction has become the bottleneck for a lot of AI applications in regulated industries, and generic OCR or single-model pipelines break on the exact edge cases (handwriting, faxes, broken tables) that these workflows depend on. The team frames Hanji around a playground you can try on real documents and a short onboarding where they tune the pipeline to a specific corpus, citing case work like lifting text accuracy from 71 percent to 96.5 percent for one healthcare customer, cutting end-to-end latency from 90 seconds to under a minute for another, and reducing processing cost by 75 percent for a third by adding a routing step. Hanji is built by Achyut Krishna Byanjankar, Soami Kapadia, and David Yu , part of Y Combinator's Spring 2025 batch. The founders previously built YouLearn, an AI tutor used by millions of students, where processing textbooks, handwritten notes, and scanned documents at scale is what pushed them to build the underlying parser that became Hanji.
Tags
<500KSeedProduct launchExplainerB2BGlobalUSVertical AIFounder-led
Visit product
Comments (7)
Sign in to join the discussion.
Priya Ranganathan15d ago

The fax reference in the copy is doing heavy lifting. Every ops team at a hospital or insurer just perked up because yes, they still receive faxes in 2025.

Tomáš K.14d ago

Going from consumer edtech to enterprise doc pipelines is a wild pivot arc. Respect for choosing the least sexy problem with the fattest wallets attached.

Kwesi O.15d ago

Nice landing page but the demo GIF cuts before showing a truly gnarly doc. Show me a coffee-stained handwritten insurance claim from 1997 or I'm not sold.

Mika Noda15d ago

Curious what your P95 latency looks like on multi-page scans with mixed handwriting. 'Pipeline tuned for accuracy, speed, and cost' is the trilemma every OCR startup claims to solve.

Rafael Vasquez14d ago

Tagline rewrite, on the house: 'Messy documents in, clean data out.' You're welcome, David.

Linnea Svärdstrom14d ago

Before anyone in regulated industries touches this: where does the data sit, who trains on it, and is there a BAA? Would love to see that surfaced above the fold.

obi14d ago

Two screenshots and one sentence for a YC launch tweet is bold. Either supreme confidence or you ran out of Canva credits, either way it worked because I clicked.