How Would You Guys Go About Cleaning Up PDF Data?

I’m trying to take the CDSs (common data sets) of a bunch of universities and compare them together, but I need to find some way to automate the process of extracting the data from them (probably into a SQL database). The issue is that although the questions on the forms are standardized, some universities convery it very differently. For example, look at C7 on the Stanford and Princeton common data sets.

So how should I go about doing this? I tried to leverage Claude’s sonnet model but it didn’t go too well, the context was too large for Claude and it was mixing up multiple fields.

And using something like tabula or pdfplumber doesn’t really help since the universities format it so differently.

Any advice would be appreciated, thank you!

submitted by /u/Roxy201
[link] [comments]

Leave a Reply

Your email address will not be published. Required fields are marked *