Skip to main content
Question

Extracting PDF Annotations and Transferring Them to Point Features

  • September 30, 2026
  • 1 reply
  • 6 views

anandkumargis
Participant
Forum|alt.badge.img+4

Hi FME Folks,

I’m working with a georeferenced PDF that contains several annotations/marked-up text. I also have an existing point shapefile/point feature class.

My requirement is to:

  • Read the georeferenced PDF in FME.
  • Extract the annotation text.
  • Get the location/coordinates of each annotation.
  • Transfer the extracted annotation text to the corresponding point feature in the shapefile.
  • Write the annotation text into an attribute field of the point feature class.

Has anyone implemented a similar workflow in FME?

I would appreciate any suggestions on the appropriate FME reader/transformers or the best approach to extract both the annotation text and its spatial location from a georeferenced PDF.

Thanks in advance for your suggestions!

Anand

 

1 reply

david_r
Celebrity
  • September 30, 2026

The FME PDF reader works fairly well, but it will depend a lot on how your PDFs are structured internally, as the PDF format isn’t meant for parsing, but for reading on a page. The issue is that FME is making a lot of assumptions on e.g. character spacing and line heights when assembling the characters into words and blocks of text. If FME can extract the text in strings that suit your use case, then it will probably be fairly easy -- you simply have to test with your PDFs and see what FME returns.

However, if FME struggles to return anything sensible from your PDF it might be more challenging, as FME does not let you specify any parameters to the internal algorithm that groups characters into strings.

I’ve done something similar where FME was unable to return anything but individual characters, and I had to resort to using Python to get anything sensible: https://pdfminersix.readthedocs.io/en/latest/ was really great for this.

If your PDFs doesn’t contain anything secret, you may also want to consider using a light-weight LLM for parsing.