Examples:
- picture n1: this is a random page from a PDF, it shows 6 different pictures. Getting all of them together as 1 is not an issue;
- picture n2: this is a non-obvious part of the top-right graph, I have no idea how to tackle this
- picture n3: this is a more obvious part of the bottom-left graph, easier to extract, but when doing so I still get the "noise text" from the graph on the right, as there is some coordinates overlap. I have no idea how to tackle this.
Not looking for the coding solution (although of course, it would help lol), but I'd highly appreciate:
- mental framework on how to divide the problem into subproblems
- any tools that I can use? (ie. unstructured.io is not great for this)
- any technologies or specific ML/computer vision libraries I can look into?
Thanks:)