Repository navigation
fix(docx): preserve underlined text - #2017
Adam Fourney (afourney) merged 8 commits into
Conversation
|
Local re-check after cleaning up my environment:
The explicit |
|
Copilot resolve the merge conflicts in this pull request |
There was a problem hiding this comment.
🟡 Changes recommended
The OCR DOCX converter still drops underlines when plugins are enabled.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Preserves DOCX underline formatting as HTML in Markdown output.
Changes:
- Adds Mammoth’s default underline style mapping.
- Preserves
<u>tags during Markdown conversion. - Adds an end-to-end DOCX regression test.
File summaries
| File | Description |
|---|---|
_docx_converter.py |
Adds underline style mapping. |
_markdownify.py |
Retains underline HTML markup. |
test_module_misc.py |
Tests underlined DOCX text. |
Review details
- Files reviewed: 3/3 changed files
- Comments generated: 2
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…ix/docx-preserve-underline # Conflicts: # packages/markitdown/src/markitdown/converters/_markdownify.py # packages/markitdown/tests/test_module_misc.py
Extract the underline style-map default into a shared with_underline_style_map() helper and use it in both of DocxConverterWithOCR's mammoth paths, so underlines survive DOCX conversion in plugin mode as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Mammoth composes style mappings as custom + embedded + built-in defaults, first match wins, so appending "u => u" to the caller's map placed it ahead of any map embedded in the document: an embedded "u => em" silently lost to the default. Compose caller + embedded + "u => u" explicitly and call mammoth with include_embedded_style_map=False, in a convert_docx_to_html() helper shared by DocxConverter and DocxConverterWithOCR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There was a problem hiding this comment.
🟡 Changes recommended
The OCR package imports a new core helper without updating its compatible markitdown dependency range.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Balanced
There was a problem hiding this comment.
🟡 Changes recommended
The OCR converter references pre_process_stream before assignment, causing every conversion to fail.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Balanced
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Summary
Fixes #35.
DOCX underlined runs were lost because Mammoth only emits underline markup when a style map for
uis present, and the HTML-to-Markdown pass did not preserve<u>tags. This adds the missing default style map entry for DOCX conversion and keeps underline markup in the Markdown output.To verify
PYTHONPATH=packages/markitdown/src python -m pytest packages/markitdown/tests/test_module_misc.py -k docx -qPYTHONPATH=packages/markitdown/src python -m pytest packages/markitdown/tests/test_module_misc.py::test_docx_underlined_text_is_preserved packages/markitdown/tests/test_module_misc.py::test_docx_comments -qpython -m py_compile packages/markitdown/src/markitdown/converters/_docx_converter.py packages/markitdown/src/markitdown/converters/_markdownify.py packages/markitdown/tests/test_module_misc.pygit diff --check