graph LR
Application_Interface_Entry_Points_["Application Interface (Entry Points)"]
Core_Processing_Engine["Core Processing Engine"]
External_Tooling_Integration["External Tooling & Integration"]
PDF_Document_Analysis["PDF Document Analysis"]
System_Utilities_Framework["System Utilities & Framework"]
Application_Interface_Entry_Points_ -- "Initiates" --> Core_Processing_Engine
Application_Interface_Entry_Points_ -- "Utilizes" --> System_Utilities_Framework
Core_Processing_Engine -- "Orchestrates" --> PDF_Document_Analysis
Core_Processing_Engine -- "Orchestrates" --> External_Tooling_Integration
Core_Processing_Engine -- "Utilizes" --> System_Utilities_Framework
Core_Processing_Engine -- "Produces output for" --> Application_Interface_Entry_Points_
External_Tooling_Integration -- "Provides services to" --> Core_Processing_Engine
External_Tooling_Integration -- "Provides services to" --> PDF_Document_Analysis
External_Tooling_Integration -- "Utilizes" --> System_Utilities_Framework
PDF_Document_Analysis -- "Provides data to" --> Core_Processing_Engine
PDF_Document_Analysis -- "Utilizes" --> External_Tooling_Integration
PDF_Document_Analysis -- "Utilizes" --> System_Utilities_Framework
System_Utilities_Framework -- "Supports" --> Application_Interface_Entry_Points_
System_Utilities_Framework -- "Supports" --> Core_Processing_Engine
System_Utilities_Framework -- "Supports" --> PDF_Document_Analysis
System_Utilities_Framework -- "Supports" --> External_Tooling_Integration
click Application_Interface_Entry_Points_ href "https://github.com/CodeBoarding/GeneratedOnBoardings/blob/main//OCRmyPDF/Application_Interface_Entry_Points_.md" "Details"
click Core_Processing_Engine href "https://github.com/CodeBoarding/GeneratedOnBoardings/blob/main//OCRmyPDF/Core_Processing_Engine.md" "Details"
click PDF_Document_Analysis href "https://github.com/CodeBoarding/GeneratedOnBoardings/blob/main//OCRmyPDF/PDF_Document_Analysis.md" "Details"
click System_Utilities_Framework href "https://github.com/CodeBoarding/GeneratedOnBoardings/blob/main//OCRmyPDF/System_Utilities_Framework.md" "Details"
One paragraph explaining the functionality which is represented by this graph. What the main flow is and what is its purpose.
This component serves as the primary user-facing layer, responsible for parsing command-line arguments or API calls, validating initial inputs, and orchestrating the initiation of the OCR processing workflow. It acts as the gateway for all user interactions with OCRmyPDF.
Related Classes/Methods:
numeric(0:0)str_to_int(0:0)ArgumentParser(0:0)LanguageSetAction(0:0)get_parser(0:0)run(0:0)get_parser_options_plugins(0:0)
This is the central orchestrator of the OCR workflow. It manages the sequence of operations including input preparation, OCR execution, HOCR generation, transformation of HOCR into a searchable text layer, and final PDF assembly and optimization. It coordinates various sub-processes to achieve the final OCR'd PDF.
Related Classes/Methods:
run_pipeline_cli(0:0)
This component provides an abstracted and standardized interface for interacting with external command-line tools (e.g., Tesseract, Ghostscript, jbig2enc). It handles subprocess execution, argument formatting, and error reporting, encapsulating the specifics of each external tool's integration.
Related Classes/Methods: None
This component specializes in extracting detailed structural and content information from PDF documents. This includes identifying pages, images, existing text layers, and analyzing their layout, which is critical for determining the necessity and strategy for OCR processing.
Related Classes/Methods: None
This is a foundational layer providing cross-cutting concerns and general-purpose functionalities. This includes input/option validation, PDF metadata management, concurrency control, job context management (temporary files, logging), general helper utilities, and the plugin management system that allows for extending OCRmyPDF's capabilities.
Related Classes/Methods:
check_options(0:0)configure_logging(0:0)