GLM-OCR
### TL;DR
GLM-OCR is a lightweight professional OCR model with parameters as small as 0.9B, yet it achieves state-of-the-art performance across multiple capabilities. It sets a new benchmark for document parsing with its "small size and high accuracy." Key features include: - Performance SOTA: Scored 94.62 points to top OmniDocBench V1.5 and achieved current best performance across multiple mainstream document understanding benchmarks including tables and formulas at launch. - Optimized for Real-World Scenarios: Delivers stable, leading accuracy in complex environments like code documentation, intricate tables, and stamp recognition. Maintains exceptional recognition precision even with complex layouts, diverse fonts, or mixed text-image content. - Efficient and Cost-Effective: With just 0.9B parameters, supports VLLM and SGLang deployment, significantly reducing inference latency and computational overhead.
Key Insights & Metrics
Key Features
- State-of-the-art performance with a score of 94.62 on OmniDocBench V1.5
- Optimized for complex environments like code documentation and intricate tables
- Efficient and cost-effective with only 0.9B parameters, supporting VLLM and SGLang deployment
Discussion
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!