#๐Ÿ”’ How to extract non-obvious sub-parts of graphs inside a PDF?

4 messages ยท Page 1 of 1 (latest)

earnest perch
#

Examples:

  • picture n1: this is a random page from a PDF, it shows 6 different pictures. Getting all of them together as 1 is not an issue;
  • picture n2: this is a non-obvious part of the top-right graph, I have no idea how to tackle this
  • picture n3: this is a more obvious part of the bottom-left graph, easier to extract, but when doing so I still get the "noise text" from the graph on the right, as there is some coordinates overlap. I have no idea how to tackle this.

Not looking for the coding solution (although of course, it would help lol), but I'd highly appreciate:

  • mental framework on how to divide the problem into subproblems
  • any tools that I can use? (ie. unstructured.io is not great for this)
  • any technologies or specific ML/computer vision libraries I can look into?

Thanks:)

late pierBOT
#

@earnest perch

Python help channel opened

Remember to:

  • Ask your Python question, not if you can ask or if there's an expert who can help.
  • Show a code sample as text (rather than a screenshot) and the error message, if you've got one.
  • Explain what you expect to happen and what actually happens.

:warning: Do not pip install anything that isn't related to your question, especially if asked to over DMs.

late pierBOT
#

@earnest perch

Python help channel closed

This help channel has been closed and it's no longer possible to send messages here. If your question wasn't answered, feel free to create a new post in #1035199133436354600. To maximize your chances of getting a response, check out this guide on asking good questions.