🔧 Công cụ lập trình•Mã nguồn mở•Đang hoạt động
Chandra
Mô hình OCR xử lý bảng biểu phức tạp, form, chữ viết tay với khả năng nhận diện layout đầy đủ.
#chandra
Danh mục
🔧 Công cụ lập trình
Giá
Mã nguồn mở
GitHub Stars
⭐ 12,384
Ngôn ngữ
Python
License
Apache-2.0
Ngày thêm
2026-03-29
Tóm tắt từ README GitHub
Datalab
State of the Art models for Document Intelligence
Chandra OCR 2
Chandra OCR 2 is a state of the art OCR model that converts images and PDFs into structured HTML/Markdown/JSON while preserving layout information.
Try Chandra on Datalab
Our managed platform runs an improved Chandra with higher accuracy than the open weights, zero data retention by default, SOC 2 Type 2, and custom BAAs.
If you have high volume workloads, we offer a batch processing service that has processed 200M+ pages per week — we manage the infrastructure so your workloads finish on time.
Get started with $5 in free credits — sign up — takes under 30 seconds — or try Chandra in our public playground.
Commercial self-hosting requires a license — see Commercial usage. For on-prem licensing, contact us.
News
- 3/2026 - Chandra 2 is here with significant improvements to math, tables, layout, and multilingual OCR
- 10/2025 - Chandra 1 launched
Features
- Tops external olmocr benchmark and significant improvement in internal multilingual benchmarks
- Convert documents to markdown, html, or json with detailed layout information
- Support for 90+ languages (benchmark below)
- Excellent handwriting
Xem thêm từ README.mdThu gọn README.md
Đánh giá chi tiết
Tổng quan
Chandra là mô hình OCR mã nguồn mở từ Datalab, chuyên xử lý tài liệu phức tạp: bảng biểu, form, chữ viết tay. Điểm mạnh là giữ nguyên cấu trúc layout gốc thay vì chỉ trích xuất text thô.
Tính năng chính
- Nhận diện bảng biểu phức tạp (merged cells, nested tables)
- Xử lý form có field labels và checkboxes
- OCR chữ viết tay với độ chính xác cao
- Output giữ nguyên layout, dễ convert sang structured data
Ai nên dùng
Developer cần OCR tài liệu phức tạp (hóa đơn, form, bảng kê), hoặc ai xây pipeline xử lý tài liệu tự động.
Hạn chế
Còn khá mới (7.6k stars), ecosystem chưa phong phú bằng các OCR engine lâu đời hơn.