CyberRota Analysis
AI-GeneratedThe Tesseract OCR engine versions 5.5.3 and earlier are vulnerable due to a flaw in the GenericVector::read callback, which allows for heap out-of-bounds writes when processing crafted .traineddata files. This vulnerability can lead to heap corruption, application crashes, or potentially allow for controlled memory corruption, posing a significant risk to systems utilizing Tesseract for OCR tasks. Organizations using Tesseract should prioritize this issue, especially those handling untrusted input data or operating in security-sensitive environments, as no fix is currently available.
Public Exploit Signal
A public exploit, PoC, GitHub repository or Metasploit reference was detected for this CVE.
Note: these links are listed for security research and verification purposes only.
Original NVD Description
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, the callback form of GenericVector::read in src/ccutil/genericvector.h reads the independent int32 fields reserved and size_used_ from a .traineddata model without a cap or an invariant check. reserve(reserved) allocates the backing array, but the callback loop writes size_used_ elements. A crafted TESSDATA_INTTEMP component with version_id 4 or later can therefore set reserved to a small value and size_used_ to a large value when fontinfo_table_.read(fp, read_info) is called from src/classify/intproto.cpp, causing a heap out-of-bounds write of FontInfo structures, heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.