This comparison of optical character recognition software includes:

  • OCR engines, that do the actual character identification
  • Layout analysis software, that divide scanned documents into zones suitable for OCR
  • Graphical interfaces to one or more OCR engines
  • Software development kits that are used to add OCR capabilities to other software (e.g. forms processing applications, document imaging management systems, e-discovery systems, records management solutions)
Sortable table
Name Founded year Latest stable version Release year License Online Windows Mac OS X Linux BSD Android iOS Programming language SDK? Languages Fonts Output Formats Notes
ABBYY FineReader1989162022ProprietaryYesYesYesNoYes Yes YesC/C++Yes192[1]All fontsDOC, DOCX, XLS, XLSX, PPTX, RTF, PDF, HTML, CSV, TXT, ODT, DjVu, EPUB, FB2[2]ABBYY also supplies SDKs for embedded and mobile devices. Professional, Corporate and Site License Editions for Windows, Express Edition for Mac.[3]
AnyDoc Software1989??ProprietaryNoYesNoNoNo ? ?VBScript???Works with structured, semi-structured, and unstructured documents.
Asprise OCR SDK1998152015ProprietaryYesYesYesYesYes ? ?Java, C#,VB.NET, C/C++/DelphiYes20+[4]?Plain text, searchable PDF, XML[5]Java, C#, VB.NET, C/C++/Delphi SDKs for OCR and Barcode recognition on Windows, Linux, Mac OS X and Unix.[6]
CuneiForm19961.12011BSD variantNoYesYesYesYes ? ?C/C++Yes28Any printed fontHTML, hOCR, native, RTF, TeX, TXT[7]Enterprise-class system, can save text formatting and recognizes complicated tables of any structure
Dynamsoft OCR SDK20038.22012ProprietaryYesYesNoNoNo ? ?C/C++Yes40+[8]?PDF, TXT
E-aksharayan 2010 Yes No Yes No ? ? 14 RTF, TXT, BRL
GOCR20000.52[9]2018GPLYes[10]YesYesYesYes ? ?C?20+?
Google Drive OCR or Google Cloud Vision2015ProprietaryYesBrowserBrowserBrowserUnknown ? ?UnknownYes200+All fontstextGoogle blog post[11][12]
Microsoft Office Document Imaging?Office 20072007ProprietaryNoYesNoNoNo ? ?????Uses OmniPage
Microsoft Office OneNote 20072011?2007ProprietaryNoYesNoNoNo ? ?????
OCRFeeder2009-030.8.52022GPLNoNoNoYesNo ? ?Python???Features a full user interface and has a command-line tool for automatic operations. Has its own segmentation algorithm but uses system-wide OCR engines like Tesseract or Ocrad
Ocrad?0.28[13]2022GPLYesNoYesYesYes ? ?C++YesLatin alphabet?Command line
OCRopus20071.3.32017ApacheNoNoYesYesYes ? ?Python?All languages using Latin script (other languages can be trained)Normal Latin script and Fraktur (other scripts can be trained)TXT, hOCR,[14] PDF[15]Pluggable framework under active development, used for Google Books
OmniPage1970s19.22015ProprietaryYesYesYesYesNo ? ?C/C++, C#[16]Yes125[17]Machine and handprinted fontsDOC/DOCX XLS/XLSX PPTX RTF PDF PDF/A Searchable PDF HTML Text XML ePUB MP3Product of Nuance Communications
Puma.NET??2009BSDNoYesNoNoNo ? ?C#Yes28Any printed font.NET OCR SDK based on Cognitive Technologies' CuneiForm recognition engine. Wraps Puma COM server and provides simplified API for .NET applications
ReadSoft???ProprietaryNoYesNoNoNo ? ?????Scan, capture and classify business documents such as invoices, forms and purchase orders integrated with business processes.
Scantron???ProprietaryNoYesNoNoNo ? ?????For working with localized interfaces, corresponding language support is required.
SmartScore199110.5.82015ProprietaryNoYesYesNoNo ? ?????For musical scores
Tesseract19855.3.32023ApacheNoYesYesYesYes ? ?C++, CYes100+[18]Any printed fontText, ALTO, hOCR,[19] PDF, others with different user interfaces[20] or the APICreated by Hewlett-Packard; under further development by Google[21]
Name Founded year Latest stable version Release year License Online Windows Mac OS X Linux BSD Android iOS Programming language SDK? Languages Fonts Output Formats Notes

Evaluation

A 2016 analysis of the accuracy and reliability of the OCR packages Google Docs OCR, Tesseract, ABBYY FineReader, and Transym, employing a dataset including 1227 images from 15 different categories concluded Google Docs OCR and ABBYY to be performing better than others.[22]

References

  1. ↑ "ABBYY FineReader 14: Technical Specifications". Finereader.abbyy.com. Retrieved 2017-02-23.
  2. ↑ "ABBYY FineReader 11: Technical Specifications". Finereader.abbyy.com. Retrieved 2013-09-12.
  3. ↑ "Top OCR Software". Ocrworld.com. 2010-03-30. Archived from the original on 2017-02-23. Retrieved 2013-09-12.
  4. ↑ "Asprise OCR SDK Features". asprise.com. Retrieved 2014-06-21.
  5. ↑ "Asprise Java OCR Library Features". asprise.com. Retrieved 2014-06-21.
  6. ↑ "Asprise Java, C#/VB.NET OCR API". asprise.com. 2015-11-19. Retrieved 2015-11-19.
  7. ↑ Debian manual page for Cuneiform for Linux version 1.1.0
  8. ↑ "OCR SDK Language Packages Download". Dynamsoft.com. Retrieved 2013-09-12.
  9. ↑ "GOCR Homepage". wasd.urz.uni-magdeburg.de. Retrieved 2018-10-17.
  10. ↑ "GOCR". Jocr.sourceforge.net. Retrieved 2013-09-12.
  11. ↑ "Supported languages". Feb 11, 2022.
  12. ↑ Ashok Popat (Sep 4, 2015). "IEEE SPS: Optical Character Recognition for Most of the World's Languages". YouTube. Archived from the original on 2021-12-20.
  13. ↑ Diaz, Antonio (2022-01-17). "GNU Ocrad 0.28 released" (Mailing list). info-gnu.
  14. ↑ OCRopus includes the ocropus-hocr tool which produces hOCR from the recognition results.
  15. ↑ In combination with the hocr-tools
  16. ↑ "OmniPage CSDK - OCR Document Capture Toolkit | Document Imaging & OCR". Nuance. Archived from the original on 2010-08-24. Retrieved 2013-09-12.
  17. ↑ "OmniPage Standard Document Conversion". Nuance. Archived from the original on 2014-03-13. Retrieved 2014-02-25.
  18. ↑ Based on count of language training files for version 3.04. Available at the download page.
  19. ↑ Usage explained in the Tesseract Readme and FAQ
  20. ↑ Such as ODF with OCRFeeder
  21. ↑ "GitHub - tesseract-ocr/tesseract: Tesseract Open Source OCR Engine (main repository)". GitHub. Retrieved 2018-11-05.
  22. ↑ Assefi, Mehdi (2016-12-01). "OCR as a Service: An Experimental Evaluation of Google Docs OCR, Tesseract, ABBYY FineReader, and Transym". ResearchGate. Retrieved 2019-01-31.
This article is issued from Wikipedia. The text is licensed under Creative Commons - Attribution - Sharealike. Additional terms may apply for the media files.