Yeah, PDF is more-or-less an open standard now: https://www.iso.org/standard/51502.html
I say “more or less” because there are still some Adobe-specific extensions on top of the base PDF version, but those are minor.
Also, Adobe Acrobat (the editing/PDF production tool, not Acrobat Reader) is still one of the better PDF editing apps out there, IMHO. It’s a nice enough standalone app, just sucks when it’s bundled into a monthly subscription.
But IMHO their real “monopoly” is in marketing and branding… same as “Microsoft” office and LibreOffice. Third-party free or cheap software can do like 90% of the same things, and most people will never need the last 10%, but people still equate Adobe with PDF and Microsoft with Office.
I think it’s a different situation for you here?
I believe what @dolphinboy is describing is hidden metadata (or some sort of otherwise useless “binary junk”) added by a poorly-behaved PDF generating app (some sort of movie editor, apparently).
I think what you’re talking about — if I understand you correctly — is people scanning you text PDFs that end up huge? In that case, it’s because the scan itself kept the text as an image, and you need to OCR it.
There’s a bunch of software out there for this, but if you have an AI chatbot you like, just upload the PDF(s) there and tell it to extract the text and make a clean PDF out of it without the scanned images. (They’re both better at OCR than most non-LLM OCR apps, and also generally easier to use and cheaper than the PDF editing apps for something like this. In a traditional PDF workflow, you’d have to OCR the text, copy the text into a separate Word doc, proofread it and fix OCR errors, reformat it, and re-save it as PDF. But an AI can do all that in one prompt, at an accuracy far far above traditional OCR software.)