A free playground from Extend, the company behind a document processing platform. Upload your own PDF, JPEG, or PNG, and two unnamed OCR/VLM models (Anonymous Model 1 and 2) read the same document so you can compare the results side by side. Vote for whichever you think did better, and your vote feeds into a public ELO-based leaderboard. Launched in November 2025, it started with more than ten models including Gemini 3, DeepSeek-OCR, and Qwen3-VL, and new models keep getting added as they arrive.
Key Features
- Anonymous battle format: Your uploaded document is processed by two anonymous models at once, letting you compare the extracted text side by side. Model names are revealed only after you vote
- Random document option: No sample on hand? Pick one at random from a prepared set and run a battle anyway
- Public ELO-based leaderboard: Votes are aggregated into a table showing rank, model name, ELO score, win rate, and battle count. Strength across different fonts, layouts, and languages becomes visible in real-world usage terms
- Open source and commercial VLMs side by side: Evaluate open source OCR models like DeepSeek-OCR against commercial multimodal models like Gemini 3 on the same footing
- Model hosting by Baseten: Model inference hosting runs through a partnership with Baseten, so Extend can focus on operating the evaluation environment
Pros and Cons
✅ Pros
- Check how models actually differ on your own documents, not just on academic benchmark numbers
- Blind voting keeps brand reputation from tilting your judgment
- No signup, no cost — try it the moment the thought crosses your mind
- The leaderboard keeps updating, making it a useful reference point whenever a new model appears
⚠️ Cons
- Leaderboard rankings shift as votes accumulate, so a screenshot from one moment won’t necessarily hold later
- Supported files are limited to PDF, JPEG, and PNG — no dedicated testing for video or handwritten documents
- No public API or documentation; it stays a browser-based comparison tool
- Extend itself offers a document processing API called “Parse,” so the selection and presentation of models under evaluation isn’t fully third-party neutral
Comparison with Similar Services
| Criterion | OCR Arena | LMArena | Academic benchmarks like OmniDocBench |
|---|---|---|---|
| What’s evaluated | OCR and VLM models (document reading focus) | Chatbots and general-purpose LLMs | OCR models (fixed dataset) |
| Evaluation method | Anonymous battles + user votes | Anonymous battles + user votes | Static scoring on quantitative metrics (accuracy, edit distance, etc.) |
| Can you use your own documents? | Yes | No (prompt-based conversation focused) | No (fixed dataset only) |
| Operator | Extend (document processing company) | LMArena team | Research institutions and companies |
Who It’s For
- Developers who want to compare reading accuracy across multiple models before adopting OCR or document analysis AI
- Teams who need to see real-world performance on formats close to their own documents, not just benchmark scores
- Anyone tracking the ongoing evolution of talked-about OCR/VLM models like DeepSeek-OCR and the Gemini family
- People gathering neutral, vendor-agnostic input for a model selection decision
Summary
OCR Arena makes the differences among today’s crowded field of OCR/VLM models visible through a mechanism anyone can grasp: anonymous battles plus a public leaderboard. Its strength is how easily you can test your own documents for free. Keep in mind that the operator is a document processing company and that leaderboard rankings are always in motion, and treat it as a first-pass reference for model selection rather than a final verdict.