# saving  .pdf file as text?

**URL:** <https://boards.straightdope.com/t/saving-pdf-file-as-text/109020>\
**Category:** Factual Questions\
**Created:** [May 14, 2002, 2:22pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020 "2002-05-14T14:22:27Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Knighted\_Vorpal\_Sword](https://avatars.discourse-cdn.com/v4/letter/k/ac91a4/32.png) [@Knighted\_Vorpal\_Sword](https://boards.straightdope.com/u/Knighted_Vorpal_Sword)\
**Post date:** [May 14, 2002, 2:22pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/1 "2002-05-14T14:22:27Z")

</div>

Is it possible to save a .pdf file as either a text or Word file? If so, how?

---

<div class="post-metadata">

**Author:** ![RealityChuck](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/realitychuck/32/195_2.png) [@RealityChuck](https://boards.straightdope.com/u/RealityChuck)\
**Post date:** [May 14, 2002, 2:24pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/2 "2002-05-14T14:24:23Z")

</div>

Not directly, but you can hit CTRL/A to select all, and then copy the text to the clipboard. It can then be pasted into a document.

---

<div class="post-metadata">

**Author:** ![handy](https://avatars.discourse-cdn.com/v4/letter/h/b5a626/32.png) [@handy](https://boards.straightdope.com/u/handy)\
**Post date:** [May 14, 2002, 2:39pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/3 "2002-05-14T14:39:40Z")

</div>

You can also use the adode website to do a free PDF to HTML & cut & paste from there if you want.  
It depends on the complexity of the doc. Just search the web for pdftohtml.

---

<div class="post-metadata">

**Author:** ![handy](https://avatars.discourse-cdn.com/v4/letter/h/b5a626/32.png) [@handy](https://boards.straightdope.com/u/handy)\
**Post date:** [May 14, 2002, 2:44pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/4 "2002-05-14T14:44:18Z")

</div>

Oh, that should be [adobe.com](http://adobe.com) 🙂

---

<div class="post-metadata">

**Author:** ![joemama24\_98](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@joemama24\_98](https://boards.straightdope.com/u/joemama24_98)\
**Post date:** [May 14, 2002, 2:57pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/5 "2002-05-14T14:57:51Z")

</div>

There’s a program called GhostView (you’ll also need GhostScript) that you can download that will perform the conversion.

---

<div class="post-metadata">

**Author:** ![11811](https://avatars.discourse-cdn.com/v4/letter/1/9d8465/32.png) [@11811](https://boards.straightdope.com/u/11811)\
**Post date:** [May 14, 2002, 5:05pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/6 "2002-05-14T17:05:53Z")

</div>

The only problem with the copy and paste is that you lose the formatting of columns, indents, etc.

It’s not without its benefits, though.

Chrome

---

<div class="post-metadata">

**Author:** ![Koxinga](https://avatars.discourse-cdn.com/v4/letter/k/4af34b/32.png) [@Koxinga](https://boards.straightdope.com/u/Koxinga)\
**Post date:** [May 14, 2002, 9:23pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/7 "2002-05-14T21:23:44Z")

</div>

I use a software package called Paper Port, which serves as a sort of driver for my Visioneer brand scanner. I first convert the PDF document into a Paper Port document, and then convert again into Word. It’s very convenient, and IIRC it preserves the column formatting.

Bear in mind a potential pitfall accomanying all of the methods mentioned in this thread: if the author of the PDF document did not perform a “paper capture” in this document, then you won’t be able to copy or convert any text at all–as far as your computer is concerned, it’s just a great big (nontext) image.

---

<div class="post-metadata">

**Author:** ![Koxinga](https://avatars.discourse-cdn.com/v4/letter/k/4af34b/32.png) [@Koxinga](https://boards.straightdope.com/u/Koxinga)\
**Post date:** [May 14, 2002, 9:25pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/8 "2002-05-14T21:25:00Z")

</div>

And of course the best way to get to the text in a PDF document is to use Adobe Acrobat, if you’re willing to spend around $300-$400 for it–you can capture your own documents and do a format-preserving copy and paste.

---

<div class="post-metadata">

**Author:** ![kanicbird](https://avatars.discourse-cdn.com/v4/letter/k/5f8ce5/32.png) [@kanicbird](https://boards.straightdope.com/u/kanicbird)\
**Post date:** [May 14, 2002, 9:51pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/9 "2002-05-14T21:51:22Z")

</div>

Paperport may be using OCR (optical character reconition) which it does use in converting scanned images to editable text - actually i’m 99% certain that’s how paperport does it because paperport stores image files not text. So your pdf file would be subject to ocr error

---

<div class="post-metadata">

**Author:** ![Koxinga](https://avatars.discourse-cdn.com/v4/letter/k/4af34b/32.png) [@Koxinga](https://boards.straightdope.com/u/Koxinga)\
**Post date:** [May 14, 2002, 10:13pm UTC](https://boards.straightdope.com/t/saving-pdf-file-as-text/109020/10 "2002-05-14T22:13:22Z")

</div>

> [@](#):
>
> \*Originally posted by k2dave \*  
> \*\*Paperport may be using OCR (optical character reconition) which it does use in converting scanned images to editable text - actually i’m 99% certain that’s how paperport does it because paperport stores image files not text. So your pdf file would be subject to ocr error \*\*

No, OCR wouldn’t enter into it because the characters have already been “recognized”, assuming that the PDF file has already been through the capture routine.
