Just in case anyone was curious, here’s what the original PDF looked like:
(black rectangles added by me to mask screenplay content)
If you decompress the PDF content, it gives you this:
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 266.4000 148.8000 Tm
(THE END) Tj % THE END, centered by starting at x=266.4. The last real line of the script
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 180.0000 136.8000 Tm
(504b030414000600080000002100e9de0fbf) Tj % JUNK BEGINS. x=180 is the dialogue margin; Screenwriter is treating the hex as dialogue under the cue above
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 180.0000 124.8000 Tm
(ff0000001c020000130000005b436f6e7465) Tj % junk, continued
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 180.0000 112.8000 Tm
(6e745f54797065735d2e786d6cac91cb4ec3) Tj % junk, continued
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 180.0000 100.8000 Tm
(301045f748fc83e52d4a9cb2400825e982c7) Tj % junk, continued
/Fcpdf2 12.000 Tf
1.0000 0.0000 -0.0000 1.0000 266.4000 88.8000 Tm
(\(MORE\)) Tj % automatic "(MORE)" cue because the dialogue runs onto the next page. Parentheses in a string must be escaped as \( and \)
/Fcpdf2 12.000 Tf
If you take all that junk output and recombine them, you get Word snippets like:
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"><Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/><Default Extension="xml" ContentType="application/xml"/><Override PartName="/theme/theme/themeManager.xml" ContentType="application/vnd.openxmlformats-officedocument.themeManager+xml"/><Override PartName="/theme/theme/theme1.xml" ContentType="application/vnd.openxmlformats-officedocument.theme+xml"/></Types>
<a:theme xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main" name="Office Theme"><a:themeElements><a:clrScheme name="Office"><a:dk1><a:sysClr val="windowText" lastClr="000000"/></a:dk1><a:lt1><a:sysClr val="window" lastClr="FFFFFF"/></a:lt1><a:dk2><a:srgbClr val="44546A"/></a:dk2><a:lt2><a:srgbClr val="E7E6E6"/></a:lt2><a:accent1><a:srgbClr val="4472C4"/></a:accent1><a:accent2><a:srgbClr val="ED7D31"/></a:accent2><a:accent3><a:srgbClr val="A5A5A5"/></a:accent3><a:accent4><a:srgbClr val="FFC000"/></a:accent4><a:accent5><a:srgbClr val="5B9BD5"/></a:accent5><a:accent6><a:srgbClr val="70AD47"/></a:accent6><a:hlink><a:srgbClr val="0563C1"/></a:hlink><a:folHlink><a:srgbClr val="954F72"/></a:folHlink></a:clrScheme><a:fontScheme name="Office"><a:majorFont><a:latin typeface="Calibri Light" panose="020F0302020204030204"/><a:ea typeface=""/><a:cs typeface=""/><a:font script="Jpan" typeface="游ゴシック Light"/><a:font script="Hang" typeface="맑은 고딕"/><a:font script="Hans" typeface="等线 Light"/><a:font script="Hant" typeface="新細� typeface="等线 Light"/><a:fo�� 고eface==""��acenace="�
So taken together with the formatting (things like a specific layout width for dialogue, the “MORE” automatically generated cues, etc.), it suggests a bug within the Screenwriter software itself. Those aren’t things any PDF engine/printer would write on its own.
It would’ve still been possible to use a regular PDF editing app to remove pages 93+, and then just edit out the bottom of page 92 manually (removing the extra text only). But in this case it was just easier to have Claude write a script to remove all the crap — because it was all hexadecimal and formatted a particular way, it was easy to isolate and remove.
This was a special case, though. Certainly it’d be nice to have a nice, general-purpose free PDF editor… does one exist? I’d like to know too. Something like LibreOffice for PDFs?
The website-based editors are good for simpler edits, but I don’t know of anything that’s as easy to use as Acrobat Pro while still having the useful features… and ideally being free, or a low one-time cost.