diff --git a/data/auto_parse/level_freeze/frozen/idx_7.jsonl b/data/auto_parse/level_freeze/frozen/idx_7.jsonl new file mode 100644 index 0000000..db4221c --- /dev/null +++ b/data/auto_parse/level_freeze/frozen/idx_7.jsonl @@ -0,0 +1,25 @@ +{"idx": 7, "order": 0, "level": 0, "span": "AMENDMENT NO.1 TO CRUDE TALL OIL AND\nBLACK LIQUOR SOAP SKIMMINGS AGREEMENT"} +{"idx": 7, "order": 1, "level": 1, "span": "This Amendment No.1 (this “Amendment”) to the Supply Agreement, dated as of March 1, 2017 (the “Effective Date”), is entered into by and between WestRock Shared Services, LLC and WestRock MWV, LLC, on behalf of the affiliates of WestRock Company (“Seller”), and Ingevity Corporation, a Delaware corporation (“Buyer”).\nWHEREAS, Seller and Buyer previously entered into that certain Crude Tall Oil and Black Liquor Soap Skimmings Agreement, effective as of January 1, 2016 (the “Agreement”; capitalized terms used herein but not defined herein shall have the meanings given them in the Agreement); and\nWHEREAS, Buyer and Seller desire to amend the Agreement as detailed below.\nNOW THEREFORE, in consideration of the mutual premises and covenants contained herein and intending to be legally bound hereby, the Parties agree that the Agreement is amended as follows:\n1.Removal of Tres Barras, Santa Catarina, Brazil Mill.\nA.    Effective on March 1, 2017 (the “Release Date”), Buyer and Seller agree to remove the Tres Barras, Santa Catarina Brazil Mill (the “Brazil Mill”) from the Agreement. The removal of the Brazil Mill from the Agreement shall not release either Party from the obligation to pay any sum that may be owing to the other Party (whether then or thereafter due) or operate to discharge any liability that had been incurred by either Party prior to such removal. Buyer and Seller agree that in order to effect the removal of the Brazil Mill from the Agreement, on the Release Date:\ni.The reference to “Tres Barras, Santa Catarina Brazil” is deleted from Section 1(B);\nii.Sections 2(D)(i) and 3(D) and Exhibit E are deleted from the Agreement and replaced with “RESERVED”;\niii.The reference to “Brazil” is deleted from Section 16;\niv.In the table contained in Exhibit B, the reference to “Tres Barras, Brazil” and the accompanying information is deleted; and\nv.The reference to “Note 3” is deleted from Exhibit B."} +{"idx": 7, "order": 2, "level": 1, "span": "B.    Brazilian Local Agreement. Seller and Buyer shall cause their respective affiliates, to sign the Mutual Termination Agreement, attached as Exhibit A hereto."} +{"idx": 7, "order": 3, "level": 1, "span": "C.    Payment. As consideration for Seller agreeing to remove the Brazil Mill from the Agreement, Buyer shall pay Seller Two Hundred Fifty Thousand Dollars ($250,000), payable within thirty (30) days of the Effective Date.\n2.Amendment to Section 1(E). Section 1(E) of the Agreement is amended by deleting such Section in its entirety and replacing it with the following:"} +{"idx": 7, "order": 4, "level": 1, "span": "E."} +{"idx": 7, "order": 5, "level": 1, "span": "Freight:"} +{"idx": 7, "order": 6, "level": 1, "span": "(ii)Buyer may request and Seller shall provide a credit for underfilled vehicles (tank trucks and/or railcars) as follows:\na.CTO:  Minimum Product weight in pounds for tank truck deliveries is 45,600.  Minimum Product weight in pounds for rail cars will be calculated as 95% of the volume capacity rating, by gallons, of the individual railcar used multiplied by 8.0 pounds/gallon.\nb.BLSS: Minimum Product weight in pounds for tank trucks is 42,750 for tank trucks originating from the Demopolis, Florence, and Panama City Mills. Minimum Product weight in pounds for tank trucks is 39,900 for tank truck originating from the Evadale Mill.\nFor example:"} +{"idx": 7, "order": 7, "level": 3, "span": "(i) Buyer is responsible for determining the mode of transportation and for providing suitable tank trucks, rail cars or barges for shipments of one hundred percent (100%) of the Products from the Mills\nAll freight charges, insurance, demurrage and all other expenses incident thereto are for Buyer’s account; provided that, if Buyer incurs third party demurrage charges due to Seller’s delay, then Seller shall reimburse Buyer for such charges.  Seller will make commercially reasonable efforts to fully load tank trucks or rail cars to minimize total cost of transportation."} +{"idx": 7, "order": 8, "level": 3, "span": "(iii) In the event that the Parties determine to: (A) ship BLSS from Mills not referenced above; (B) shipment BLSS by rail car, or (C) ship Products by barge, then the Parties shall mutually agree in writing on the minimum weight for such shipment mode."} +{"idx": 7, "order": 9, "level": 3, "span": "(iv) If the Product weight as listed on Seller’s invoices for either tank trucks or rail cars (“Actual Weight”) is less than the applicable minimum weight indicated above (the “Minimum Weight”), Buyer may request and Seller shall provide a credit equal to (Buyer’s actual freight cost for such shipment divided by the Minimum Weight) multiplied by the (the Minimum Weight – the Actual Weight)\nIf Buyer disagrees with Seller’s calculation of the Actual Weight, Seller shall provide Buyer with the opportunity to inspect the measuring process and equipment, to verify the disparity. This credit calculation report will be generated by Buyer and sent to Seller on an excel spreadsheet substantially in the form of Exhibit K hereto at the time of the other calculations for amounts payable pursuant to Exhibit G of this Agreement, or credit shall be deemed waived. Seller will provide the applicable credit as a credit memo to Buyer for use within thirty (30) days from receipt of such report against applicable invoices from Seller (or, if the Agreement has terminated, will reimburse Buyer), as provided herein."} +{"idx": 7, "order": 10, "level": 1, "span": "1.  CTO tank truck shipment with an Actual Weight of 41,700 lbs\nMinimum Weight is 45,600. Seller will provide Buyer with a credit as follows:  $1,200 actual freight cost / 45,600 lbs. * (45,600 lbs-41,700 lbs.) = $102.63."} +{"idx": 7, "order": 11, "level": 1, "span": "2.  BLSS tank truck shipment from Florence mill with an Actual Weight of 37,750 lbs. Minimum Weight is 42,750 lbs. Seller will provide Buyer with a credit as follows: $1,500 actual freight cost / 42,750 lbs. * (42,750 lbs. – 37,750 lbs.) = $175.44. 3.  CTO rail car # GATX032026, which is rated for 20,460 gallons capacity, with an Actual Weight of 150,056 lbs. Minimum Weight is 155,496 lbs. (20,460 gallons * 8.0 lbs./gallon*.95) Seller will provide Buyer with a credit as follows: $5,000 actual freight cost/ 155,496 lbs. * (155,496 lbs-150,056 lbs.) = $174.92. (v) Buyer and Seller will work in good faith to enable transportation by barge as is appropriate and mutually agreed.  The initial cost to develop and construct infrastructure for barge shipments shall be borne by Buyer and the maintenance costs for such infrastructure shall be as agreed in writing. 3.    Continuing Effect. Except as expressly modified herein, all other terms and conditions of the Agreement will remain in full force and effect. All references in the Agreement to “the Agreement” or “this Agreement” shall be deemed a reference to the Agreement as amended by this Amendment. Furthermore, this Amendment may be executed in any number of counterparts, each of which shall be deemed to be an original and all of which together shall be deemed to be one and the same instrument. A signature sent by electronic or facsimile transmission shall be as valid and binding upon the Party as an original signature of such Party."} +{"idx": 7, "order": 12, "level": 1, "span": "IN WITNESS WHEREOF, the Parties hereto have caused this Amendment to be executed by their duly authorized representatives as of the Effective Date."} +{"idx": 7, "order": 13, "level": 2, "span": "INGEVITY CORPORATION"} +{"idx": 7, "order": 14, "level": 2, "span": "WESTROCK SHARED"} +{"idx": 7, "order": 15, "level": 2, "span": "SERVICES, LLC"} +{"idx": 7, "order": 16, "level": 2, "span": "By:_/S/ S. Edward Woodcock, Jr.____"} +{"idx": 7, "order": 17, "level": 2, "span": "By:_/S/ John D. Stakel_"} +{"idx": 7, "order": 18, "level": 2, "span": "Name: S. Edward Woodcock, Jr."} +{"idx": 7, "order": 19, "level": 2, "span": "Name: John D. Stakel"} +{"idx": 7, "order": 20, "level": 2, "span": "Title: EVP & President, Performance Materials"} +{"idx": 7, "order": 21, "level": 2, "span": "Title: Senior Vice President"} +{"idx": 7, "order": 22, "level": 2, "span": "Date: March 1, 2017"} +{"idx": 7, "order": 23, "level": 2, "span": "Date: March 8, 2017"} +{"idx": 7, "order": 24, "level": 2, "span": "WESTROCK MWV, LLC"} diff --git a/data/auto_parse/level_freeze/state.json b/data/auto_parse/level_freeze/state.json index c4c4ab4..ea9b284 100644 --- a/data/auto_parse/level_freeze/state.json +++ b/data/auto_parse/level_freeze/state.json @@ -7,7 +7,8 @@ 3, 4, 5, - 6 + 6, + 7 ], "history": [ { @@ -164,6 +165,12 @@ "action": "freeze", "idx": 6, "n_records": 69 + }, + { + "ts": "2026-05-17T07:10:24", + "action": "freeze", + "idx": 7, + "n_records": 25 } ] } diff --git a/scripts/parse_doc2dict_with_config.py b/scripts/parse_doc2dict_with_config.py index 96903c4..66f9182 100644 --- a/scripts/parse_doc2dict_with_config.py +++ b/scripts/parse_doc2dict_with_config.py @@ -2314,6 +2314,175 @@ def _load_sot_span_clean(idx: int) -> str | None: return None +# Section/structural-header patterns that DISQUALIFY a record from being +# treated as a title-continuation line. A line with any of these patterns +# is part of the body, not the title. +_TITLE_CONTINUATION_DISQUALIFIERS = re.compile( + r"^\s*(?:" + r"ARTICLE\s+\w+" # "ARTICLE I", "Article 1" + r"|SECTION\s+\d" # "SECTION 1", "Section 2" + r"|EXHIBIT\s+[A-Z0-9]" # "EXHIBIT A", "EXHIBIT 10.25" + r"|SCHEDULE\s+[A-Z0-9]" # "SCHEDULE 1" + r"|APPENDIX\s+[A-Z0-9]" # "APPENDIX A" + r"|ANNEX\s+[A-Z0-9]" # "ANNEX I" + r"|WITNESSETH" # "WITNESSETH THAT:" + r"|WHEREAS" # recitals + r"|RECITALS" + r"|NOW,?\s+THEREFORE" # "NOW, THEREFORE" + r"|IN\s+WITNESS\s+WHEREOF" # signature operating clause + r"|\d+\.\s" # numbered sections "1. ..." + r"|\([a-zA-Z0-9]+\)" # lettered/roman markers "(a)", "(i)" + r"|\[[^\]]*\]" # placeholder "[***]" + r"|/s/" # signature line + r"|By:\s|Name:\s|Title:\s|Date:\s" # signature labels + r")", + re.IGNORECASE, +) + + +def _looks_like_title_continuation(title: str) -> bool: + """Return True if `title` looks like an upper-line of a multi-line + agreement title (not a section heading, not a body fragment). + + A title-continuation line: + - Has alphabetic content (at least one word of 2+ letters). + - Is predominantly uppercase letters (≥ 60% of alphabetic chars + are uppercase) — agreement titles are typeset in ALL CAPS. + - Doesn't match any structural-header / section / recital / + signature pattern. + - Doesn't end with a sentence-terminator (`.`, `:`, `;`, `?`, `!`) + — agreement titles usually break mid-phrase across lines, often + ending with conjunctions ("AND", "OF") or nouns. + """ + if not title: + return False + stripped = title.strip() + if not stripped: + return False + # Must not match any structural / section / recital / signature pattern. + if _TITLE_CONTINUATION_DISQUALIFIERS.match(stripped): + return False + # Must have at least one alphabetic word of 2+ letters. + if not re.search(r"[A-Za-z]{2,}", stripped): + return False + # Compute uppercase-ratio of alphabetic chars (>= 60% uppercase). + alpha = [c for c in stripped if c.isalpha()] + if not alpha: + return False + upper_ratio = sum(1 for c in alpha if c.isupper()) / len(alpha) + if upper_ratio < 0.60: + return False + # Must not end with a sentence-terminating punctuation. Title lines + # are noun-phrases that wrap visually; bodies end with periods. + if stripped.endswith(('.', ':', ';', '?', '!')): + return False + return True + + +def _merge_multiline_l0_title(rows: list[dict[str, Any]]) -> list[dict[str, Any]]: + """Merge a multi-line agreement title into the single L0 record. + + Some EX-10 filings typeset the agreement title across two or more + visual lines, e.g.: + + AMENDMENT NO.1 TO CRUDE TALL OIL AND + BLACK LIQUOR SOAP SKIMMINGS AGREEMENT + + doc2dict's HTML walker turns each line into its own predicted-header + section, both as siblings under the SEC `` envelope. The + last line carries the AGREEMENT/PLAN keyword and gets remapped to + depth=0 (the L0 title). The preceding line(s) are stranded as + siblings at depth=1 — but they are PART of the title, not a + standalone L1 clause. + + This post-processor: + 1. Finds the L0 record (depth=0, not envelope, scope=agreement). + 2. Walks preceding siblings (same parent_node_id, smaller + node_id, no other sibling between them) in document order. + 3. Collects consecutive preceding siblings that match the + shape of a title-continuation line — uppercase, no body, + no structural-header pattern, no sentence-terminator. + 4. Prepends each continuation title to the L0 title (preserving + source order) and marks the preceding records as envelope so + they drop from JSONL but stay in parquet for audit. + + Runs BEFORE `_split_l0_title_from_preamble` so the L0 title is + complete before the preamble split. + + Shape-based detection — no phrase blocklists. The disqualifier + regex blocks structural patterns (ARTICLE, SECTION, EXHIBIT, + WHEREAS, etc.) so legitimate body lines that follow the title + aren't absorbed. + """ + l0: dict[str, Any] | None = None + for r in rows: + if ( + r.get("depth") == 0 + and not r.get("is_envelope") + and r.get("scope") == "agreement" + ): + l0 = r + break + if l0 is None: + return rows + + l0_node_id = l0["node_id"] + l0_parent_id = l0.get("parent_node_id") + if l0_parent_id is None: + return rows + + # Find preceding siblings — same parent, smaller node_id — in + # node_id order (which is doc2dict's source order). + siblings = sorted( + (r for r in rows if r.get("parent_node_id") == l0_parent_id), + key=lambda r: r["node_id"], + ) + # Identify the L0's index among siblings. + try: + l0_pos = next(i for i, s in enumerate(siblings) if s["node_id"] == l0_node_id) + except StopIteration: + return rows + if l0_pos == 0: + return rows # L0 is the first sibling — no continuation lines. + + # Walk preceding siblings from L0 backwards. Collect a contiguous + # run of title-continuation candidates ending immediately before L0. + continuation: list[dict[str, Any]] = [] + for i in range(l0_pos - 1, -1, -1): + sib = siblings[i] + if sib.get("is_envelope") or sib.get("scope") == "trailer": + # Hit an envelope (SEC wrapper) or already-dropped record. + # Stop the walk — anything earlier than this is filing chrome, + # not title. + break + if (sib.get("cls") or "") != "predicted header": + break + if (sib.get("body_direct") or "").strip(): + break + title = (sib.get("title") or "").strip() + if not _looks_like_title_continuation(title): + break + continuation.append(sib) + + if not continuation: + return rows + + # Build combined title in source order: earliest continuation line + # first, then L0's existing title. + continuation.reverse() # source order (oldest first) + continuation_titles = [(c.get("title") or "").strip() for c in continuation] + existing_l0_title = (l0.get("title") or "").strip() + combined = "\n".join(continuation_titles + [existing_l0_title]) + l0["title"] = combined + + # Mark the continuation records as envelope so they don't double-emit + # to JSONL; parquet still records them. + for c in continuation: + c["is_envelope"] = True + + return rows + + def _split_l0_title_from_preamble(rows: list[dict[str, Any]]) -> list[dict[str, Any]]: """Split the L0 (agreement title) record from its preamble body. @@ -3086,6 +3255,94 @@ def _drop_pre_title_cover_records(rows: list[dict[str, Any]]) -> list[dict[str, return rows +def _strip_page_footer_exhibit_titles(rows: list[dict[str, Any]]) -> list[dict[str, Any]]: + """Strip bare 'Exhibit N.M' titles from non-envelope, non-subdoc + exhibit-class records that carry substantive body content. + + Some EX-10 filings stamp the exhibit number on every page as a + visual header (`Exhibit 10.7` repeated at the top of each page). + doc2dict's HTML walker promotes each repeating page-header tag + into its own `cls=exhibit` section. The FIRST one is the legitimate + SEC envelope (already flagged `is_envelope=True`); subsequent ones + are spurious — they carry substantive body text from the page they + head, but their title is just the bare exhibit identifier. + + When that bare-identifier title flows into the source-position + sorter (``_sort_records_by_source_position``) as the highest- + priority probe, it matches the FIRST occurrence of "Exhibit 10.7" + in the source, dragging the record to the document start. The + title-as-root rule then drops it as pre-title content — losing + the substantive body forever. + + The structural fix: clear the title on these records so the sorter + falls back to the body probe (which IS unique to the page's + content). The body emits as its own L1 record at the correct + source position. + + Detection (purely structural): + - cls is in `_SUBDOC_CLASSES` (exhibit/schedule/appendix/annex). + - is_envelope is False (not the SEC wrapper). + - scope is "agreement" (not already a trailer). + - title matches the bare-identifier pattern + (`^EXHIBIT|SCHEDULE|APPENDIX|ANNEX $` with no descriptive + text after a separator). + - body_direct has substantive content (non-empty after strip). + - Record is NOT a real subdoc (per ``_is_real_subdoc_title``). + + Runs AFTER ``_consolidate_real_subdocs`` so real subdoc headers + keep their identifiers, and BEFORE the position sorter so the + title-strip takes effect on sorting. + """ + if not rows: + return rows + + # Build children index for the real-subdoc test. + children_of: dict[int | None, list[dict[str, Any]]] = {} + for r in rows: + children_of.setdefault(r.get("parent_node_id"), []).append(r) + for pid in children_of: + children_of[pid].sort(key=lambda r: r["node_id"]) + + for r in rows: + cls = (r.get("cls") or "") + if cls not in _SUBDOC_CLASSES: + continue + if r.get("is_envelope"): + continue + if r.get("scope") == "trailer": + continue + title = (r.get("title") or "").strip() + body = (r.get("body_direct") or "").strip() + if not title or not body: + continue + # Bare-identifier title (no descriptive separator). + if not _BARE_SUBDOC_ID_RE.match(title): + continue + # Body must be SUBSTANTIVE — not just page chrome (whitespace, + # short page-reference like "Ex. B-98"). Use 60-char threshold: + # genuine content is paragraph-shaped, page chrome is short. + # This protects records like idx=5's "Schedule 1.1" + " \n \nEx. + # B-98" body from being stripped (the body is page footer, not + # substantive content from the page). + if len(body) < 60: + continue + # Confirm not a real subdoc — check the descriptive-child rescue + # would not apply. A real subdoc with a descriptive child gets + # its title combined in `_consolidate_real_subdocs`; if it still + # has a bare identifier here, the rescue did not fire, so it's + # safe to treat as a page-footer artifact. + kids = children_of.get(r["node_id"], []) + first_child_title = kids[0].get("title") if kids else None + if _is_real_subdoc_title( + r.get("title"), r.get("cls"), r.get("is_envelope", False), first_child_title + ): + continue + # Strip the title; the sorter will fall back to the body probe. + r["title"] = "" + + return rows + + def _drop_pre_title_position_records(rows: list[dict[str, Any]]) -> list[dict[str, Any]]: """Drop in-scope records that appear BEFORE the L0 agreement title in sorted document-order. @@ -4082,6 +4339,15 @@ def parse_one(idx: int, raw: str) -> tuple[dict[str, Any], list[dict[str, Any]]] # follow a numbered sibling into that numbered section so the tree # depths reflect the actual legal hierarchy. sections = _reparent_lettered_subsections_to_numbered_siblings(sections) + # Merge multi-line agreement titles before splitting the preamble. + # When the title is typeset across two visual lines (e.g. "AMENDMENT + # NO.1 TO CRUDE TALL OIL AND\nBLACK LIQUOR SOAP SKIMMINGS AGREEMENT"), + # doc2dict emits each line as its own predicted-header sibling. Only + # the last line carries the AGREEMENT keyword and gets remapped to + # depth=0; the upper line(s) sit at depth=1. This pass merges those + # upper-line continuations into the L0 title before any downstream + # step uses the title for sorting or scope decisions. + sections = _merge_multiline_l0_title(sections) sections = _split_l0_title_from_preamble(sections) # Cover-preamble rescue: when the L0 title appears TWICE in the # source (once on the cover page, once at the agreement start), the @@ -4132,6 +4398,15 @@ def parse_one(idx: int, raw: str) -> tuple[dict[str, Any], list[dict[str, Any]]] # - Signature-page banner chrome ("SIGNATURE PAGE TO FOLLOW", # "[Signature Page Follows]") is dropped. sections = _explode_signature_block_lines(sections) + # Strip bare "Exhibit N.M" titles from spurious page-header exhibit + # records (substantive body, not envelope, not real subdoc). The + # bare-identifier title would otherwise become the highest-priority + # sort probe and match the FIRST source occurrence (the cover-page + # header), dragging the record to the document start where the + # title-as-root rule would drop it. Clearing the title makes the + # sorter fall back to the body probe, locating the record at its + # true source position. + sections = _strip_page_footer_exhibit_titles(sections) sections = _sort_records_by_source_position(sections, idx) # Anchor synthetic subdoc body records to sit right after their # parent subdoc title (Defect 4). The position sorter may