Dear @linux and @academicchatter folks:
Please suggest libre/open source tools that allow for the extraction of text and images from scientific pdf documents?
The first tool I can think of is LibreOffice Draw
Maybe there are other tools, but I think LibreOffice Draw do the job pretty well
Edit: If the PDF has written text, you may wanna use an OCR tool, but I don’t have any to suggest
@ajayiyer@mastodon.social OCRmyPDF is exactly what you are looking for
Try Zotero. It is a complete literature databas but it’s PDF reader is very good at extracting images and text. Works on all OS, web and mobile. Native Linux client has been very smooth for me. Oh, terminal it doesn’t do though. If you want to extract a large amount in an automated way, its probably not the right tool.