Repository navigation
Fix OCR PDF converter: preserve inter-word spacing on image pages - #2666
Open
Wu Shuwen (dajiaohuang) wants to merge 1 commit into
Open
Wu Shuwen (dajiaohuang) wants to merge 1 commit into
Wu Shuwen (dajiaohuang) wants to merge 1 commit into
Conversation
Fixes microsoft#2565. When a PDF page contained embedded images, the OCR path rebuilt each text line by iterating page.chars (sorted by y then x) and joining with an empty string. That glued adjacent columns together (e.g. 'Customer Name' and 'Vendor Name Ltd' became 'Customer NameVendor Name Ltd') because chars from neighboring columns that happened to share a y coordinate were concatenated without any space between them. Use pdfplumber.page.Page.extract_text_lines(strip=False) for line reconstruction instead, which applies pdfplumber's own word/column spacing logic (the same logic used by page.extract_text()). Keep the existing extract_text() fallback if extract_text_lines is unavailable or returns nothing. Update the complex-layout test fixture expectation: it was asserting the buggy glued form ('ItemQuantity', 'Widget A5') and now expects the correctly spaced output ('Item Quantity', 'Widget A 5').
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2565.
When a PDF page contained embedded images, the OCR path rebuilt each text line by iterating
page.chars(sorted by y then x) and joining with an empty string. That glued adjacent columns together (e.g. 'Customer Name' and 'Vendor Name Ltd' became 'Customer NameVendor Name Ltd') because chars from neighboring columns that happened to share a y coordinate were concatenated without any space between them.Use
pdfplumber.page.Page.extract_text_lines(strip=False)for line reconstruction instead, which applies pdfplumber's own word/column spacing logic (the same logic used bypage.extract_text()). Keep the existingextract_text()fallback ifextract_text_linesis unavailable or returns nothing. Update the complex-layout test fixture expectation: it was asserting the buggy glued form ('ItemQuantity', 'Widget A5') and now expects the correctly spaced output ('Item Quantity', 'Widget A 5').