Skip to content

Fix OCR PDF converter: preserve inter-word spacing on image pages - #2666

Open
Wu Shuwen (dajiaohuang) wants to merge 1 commit into
microsoft:mainfrom
dajiaohuang:fix/ocr-pdf-column-spacing
Open

Wu Shuwen (dajiaohuang) wants to merge 1 commit into
microsoft:mainfrom
dajiaohuang:fix/ocr-pdf-column-spacing

Conversation

@dajiaohuang

Copy link
Copy Markdown

Fixes #2565.

When a PDF page contained embedded images, the OCR path rebuilt each text line by iterating page.chars (sorted by y then x) and joining with an empty string. That glued adjacent columns together (e.g. 'Customer Name' and 'Vendor Name Ltd' became 'Customer NameVendor Name Ltd') because chars from neighboring columns that happened to share a y coordinate were concatenated without any space between them.

Use pdfplumber.page.Page.extract_text_lines(strip=False) for line reconstruction instead, which applies pdfplumber's own word/column spacing logic (the same logic used by page.extract_text()). Keep the existing extract_text() fallback if extract_text_lines is unavailable or returns nothing. Update the complex-layout test fixture expectation: it was asserting the buggy glued form ('ItemQuantity', 'Widget A5') and now expects the correctly spaced output ('Item Quantity', 'Widget A 5').

Fixes microsoft#2565.

When a PDF page contained embedded images, the OCR path rebuilt each
text line by iterating page.chars (sorted by y then x) and joining with
an empty string. That glued adjacent columns together (e.g.
'Customer Name' and 'Vendor Name Ltd' became 'Customer NameVendor Name
Ltd') because chars from neighboring columns that happened to share a y
coordinate were concatenated without any space between them.

Use pdfplumber.page.Page.extract_text_lines(strip=False) for line
reconstruction instead, which applies pdfplumber's own word/column
spacing logic (the same logic used by page.extract_text()). Keep the
existing extract_text() fallback if extract_text_lines is unavailable
or returns nothing. Update the complex-layout test fixture expectation:
it was asserting the buggy glued form ('ItemQuantity', 'Widget A5')
and now expects the correctly spaced output ('Item Quantity',
'Widget A 5').
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

markitdown-ocr: columns are glued together on PDF pages that contain an embedded image

1 participant