LINKED LIST [txt mode] ▸ What's so hard about PDF text extract...
home explore | log in

What's so hard about PDF text extraction?

filingdb.com · first added by @afreshcup · 2026-09-30 · 1 upvotes

log in to save, upvote or flag this.


─── In 0 lists ─────────────────────────────────────────

(not in any lists yet)


─── Discussions ────────────────────────────────────────

* What's so hard about PDF text extraction?
406 pts · 235 comments · node
* What's so hard about PDF text extraction?
11 pts · 15 comments · node
* What's so hard about PDF text extraction?
733 pts · 342 comments · node
see all 4
* What's so hard about PDF text extraction?
3 pts · 1 comment · node

─── From the discussion ────────────────────────────────

* GitHub - camelot-dev/camelot: A Python library to extract tabular data from PDFs
github.com · node
* GitHub - coolwanglu/pdf2htmlEX: Convert PDF to HTML without losing text or format.
github.com · node
* The sad state of PDF-Accessibility of LaTex Documents
umij.wordpress.com · node
* Peter Selinger: Creating high-quality PDF/A documents using LaTeX
mathstat.dal.ca · node
* Why GOV.UK content should be published in HTML and not PDF
gds.blog.gov.uk · node

─── Related hn threads ─────────────────────────────────

* The sad state of PDF-Accessibility of LaTex Documents (2016)
79 pts · 74 comments · node
* PDF processing and analysis with open-source tools (2021)
186 pts · 40 comments · node
* Examples to compare OCR services: Amazon vs. Google vs. Microsoft
262 pts · 65 comments · node