Improve table detection and header refinement in find_tables() - #5136
Merged
Merged
Conversation
veget-able
force-pushed
the
table-round2-rules-admission
branch
from
September 24, 2026 00:29
c47ffec to
138a442
Compare
Contributor
Author
|
The two Ubuntu jobs fail in tests/test_widget_reproducers.py (the widget reproducer tests added to main on 23 September), not in the table tests, which all pass. The same tests fail on a plain 1.28.2 install here as well, so they look independent of this PR. Happy to rebase again once those are addressed on main. |
JorjMcKie
self-requested a review
September 24, 2026 14:00
JorjMcKie
approved these changes
Sep 24, 2026
…picture candidates, split stacked tables find_tables() changes, all behind the existing opt-in keywords: - union=True forwards add_lines / add_boxes into the nested line finder; they were silently dropped before. - lines_strict: thin rectangles batched into one large fill-only path are kept as rule material when they are line-like and longer than the minimum edge length (some producers draw every grid rule as a filled rectangle inside a single path). The path's own bbox contributes no envelope, so dot screens and glyph fragments do not become tables. - union admission runs on the finder's cell groups (TableFinder(cell_group_filter=)): a candidate lying mostly inside a layout "picture" group is admitted only with table-shaped evidence (two-dimensional text support, at least two rows and two columns each at least half populated, mostly rectangular cell coverage), so a chart's gridlines are not turned into a table while a sparse matrix table is. - refine=True splits a grid whose confirmed leading header row recurs verbatim into one table per stack, reusing the placement grid already built for the header/body boundary. - Table.header is computed on first access (same value, restored extraction flags); drawings are read once per find_tables() call and shared with the refine helpers; word-in-rect selection uses a per-page centre index. Default find_tables() output is byte-identical on the test corpus (default, lines_strict and refine=True paths). Based on the review round by Minsu Kim.
find_tables(refine=True) decides the header rows from the placement grid's text alone, so a header cell spanning several (but not all) columns can end the header one row early and leave the row that names those columns in the body. After the existing header rules, the header now extends over the next row when that row supplies a single-column label under every column of such a span, puts no plain number under a column that is already labeled, is not one full-width cell, and does not repeat the blank/number/text shape of the row after it (the first record of a regular body). Nested spans repeat the step; the header never shrinks. The rule reads only the placement grid's text and spans and runs where the refined placements are tagged, so it can change only the th/td tags, to_html() and Table.header_rows of refined tables. The default, lines_strict and refine=False paths do not reach it. Based on the header rules by Minsu Kim.
veget-able
force-pushed
the
table-round2-rules-admission
branch
from
September 26, 2026 13:51
138a442 to
4572a91
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This follows #5057 and addresses cases where
find_tables()misses tables, detects chart grids as tables, or merges separate tables. It also improves header tagging for spanning cells and reduces repeated work during extraction.add_linesandadd_boxesto the nested line finder whenunion=True, so they participate in table detection.lines_strict, retain thin filled rectangles longer thanedge_min_lengtheven when many are grouped into one fill path. The path's overall bounding box is excluded from the grid. This is the largest individual gain on ParseBench, concentrated in one document family that draws tables this way.Table. Candidates mostly inside a layoutpictureregion must have text distributed across at least two rows and two columns, each at least half populated, with mostly rectangular cell coverage. This helps reject chart grids while retaining tables misclassified as pictures.refine=True, split a grid when a confirmed leading header row repeats verbatim. The header must match the column structure, and each resulting table must retain at least one row below the header.refine=True, extend the header when the next row supplies individual column labels beneath a cell spanning several, but not all, columns and does not match the body pattern. At least one body row remains. This rule changesheader_rowsand<th>/<td>tagging into_html()without changing bounding boxes orextract()output.Table.headeron first access using the original detection call's extraction flags. Reuse vector drawings within eachfind_tables()call and index word centres per page to speed up word selection.On the 1,463-page regression corpus, output is byte-for-byte unchanged in both the default and
lines_strictconfigurations. Withrefine=True, one page changes because its column-label row is now included in the header. This PR adds 16 tests intests/test_tables.py.On the ParseBench table group (503 pages,
pymupdf4llm_markdownpipeline), this PR together with the companion pymupdf4llm reading-order change raises GTRM from 72.00 on PyMuPDF 1.28.2 to 75.64. The header-completion rule contributes 0.58 points, with 10 pages improving and none regressing. For the other changes, 34 pages improve and two score lower. Those two cases are a chart grid that previously received partial credit and a datasheet header block inside an image, both now rejected by the picture-region filter.Output on the held-out DP-Bench set (200 pages, 55 tables, evaluated with TEDS) is unchanged from 1.28.2. With both PRs applied, total wall time on the 503-page ParseBench set is 13% lower than 1.28.2. Processing the additional thin rectangles can add up to 0.85 seconds on individual pages containing thousands of them.