# Maddening problem with downloading text

**URL:** <https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389>\
**Category:** Factual Questions\
**Created:** [December 22, 2016, 6:18pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389 "2016-12-22T18:18:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![dougie\_monty](https://avatars.discourse-cdn.com/v4/letter/d/439d5e/32.png) [@dougie\_monty](https://boards.straightdope.com/u/dougie_monty)\
**Post date:** [December 22, 2016, 6:18pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/1 "2016-12-22T18:18:40Z")

</div>

Every now and then I need to download some text from a website; most often it’s legal text. This time it was a California ballot measure that passed in the last election.  
To save space (and paper) I copy it into a Word format. The problem is a maddening symbol that appears when I click on the “paragraph” icon on the toolbar: a “broken arrow” that the nitwit who formatted the document when uploading it, put at the end of _EACH_ line in the document. I want to put it onto a Word document, but it has 276 pages, about half of which is empty space. I don’t have that “broken arrow” symbol in my working font, so I can’t simply use the Replace function to delete the “broken arrow” from the document and give the text normal margins! Is there any way around this? :mad:

---

<div class="post-metadata">

**Author:** ![voltaire](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/voltaire/32/313_2.png) [@voltaire](https://boards.straightdope.com/u/voltaire)\
**Post date:** [December 22, 2016, 6:21pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/2 "2016-12-22T18:21:42Z")

</div>

You’re just copying and pasting the text from the website into a Word document?

---

<div class="post-metadata">

**Author:** ![voltaire](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/voltaire/32/313_2.png) [@voltaire](https://boards.straightdope.com/u/voltaire)\
**Post date:** [December 22, 2016, 6:23pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/3 "2016-12-22T18:23:43Z")

</div>

If so, try pasting it into plain .txt with Notepad or WordPad as an interim step.

---

<div class="post-metadata">

**Author:** ![beowulff](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/beowulff/32/542_2.png) [@beowulff](https://boards.straightdope.com/u/beowulff)\
**Post date:** [December 22, 2016, 6:24pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/4 "2016-12-22T18:24:43Z")

</div>

Well, you are using the wrong tool.  
Use a text editor, not Word. On the Mac, I’d suggest Text Wrangler.

The “broken arrow” is probably are hard return - you could try selecting from the end of one line to the beginning of the next, and pasting into the search & replace field.

---

<div class="post-metadata">

**Author:** ![dracoi](https://avatars.discourse-cdn.com/v4/letter/d/90db22/32.png) [@dracoi](https://boards.straightdope.com/u/dracoi)\
**Post date:** [December 22, 2016, 6:29pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/5 "2016-12-22T18:29:10Z")

</div>

My guess is that you’re seeing a non-printing character for a line break. You’d have to check your MS Word view options to be sure it is displaying that character. If I’m right, you can access that character in a Find/Replace by typing ^l into Word’s search box.

These manual line breaks would not likely be the fault of the web programmer, but of how the copy/paste works between Word and your browser. I’ve seen this before. It’s common with PDF files as well.

As an alternate solution, you might try “Save As” from your web browser. Save the html file and open that with Word.

---

<div class="post-metadata">

**Author:** ![dougie\_monty](https://avatars.discourse-cdn.com/v4/letter/d/439d5e/32.png) [@dougie\_monty](https://boards.straightdope.com/u/dougie_monty)\
**Post date:** [December 22, 2016, 7:16pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/6 "2016-12-22T19:16:26Z")

</div>

I used Notepad to copy the document. Then I transferred it to regular Word, and did the formatting from there. (To save the paragraph demarcations I did want I keyed in “XX” at each one. Then, I used Replace to remove all those Paragraph symbols; then I replaced the “XX” marks with new Paragraph symbols. Best formatting idea you (and I) ever came up with. 🙂 )

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [December 22, 2016, 7:33pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/7 "2016-12-22T19:33:58Z")

</div>

There is a simpler way to do that, assuming that there is a hard line return between paragraphs along with the one return after each line. Just do a find/replace for all double returns (^l^l) and replace them with a paragraph mark (^p), then replace all the single returns that are left with a space mark.

---

<div class="post-metadata">

**Author:** ![dougie\_monty](https://avatars.discourse-cdn.com/v4/letter/d/439d5e/32.png) [@dougie\_monty](https://boards.straightdope.com/u/dougie_monty)\
**Post date:** [December 22, 2016, 7:41pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/8 "2016-12-22T19:41:08Z")

</div>

> [@Darren\_Garrison](#):
>
> There is a simpler way to do that, assuming that there is a hard line return between paragraphs along with the one return after each line. Just do a find/replace for all double returns (^l^l) and replace them with a paragraph mark (^p), then replace all the single returns that are left with a space mark.

Is that an upper-class “I” or lower-case “L”?

---

<div class="post-metadata">

**Author:** ![yearofglad](https://avatars.discourse-cdn.com/v4/letter/y/b782af/32.png) [@yearofglad](https://boards.straightdope.com/u/yearofglad)\
**Post date:** [December 22, 2016, 7:50pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/9 "2016-12-22T19:50:31Z")

</div>

> [@dougie\_monty](#):
>
> Is that an upper-class “I” or lower-case “L”?

They’re lower-case Ls. (I found out by copying and pasting into Word, then typing an upper I and a lower L and changing fonts until the two characters were distinctly different, then comparing what I typed to what I cut-and-pasted.)

---

<div class="post-metadata">

**Author:** ![Saintly\_Loser](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/saintly_loser/32/4045_2.png) [@Saintly\_Loser](https://boards.straightdope.com/u/Saintly_Loser)\
**Post date:** [December 22, 2016, 8:35pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/10 "2016-12-22T20:35:28Z")

</div>

Try this:

Advanced search (CTRL-H). Search for two line breaks (^l^l), replace with paragraph break (^p). That’s assuming they’re putting two lines between paragraphs.

Then search for line breaks (^l) and replace with space (just hit the space bar once).

That should give you text broken up into paragraphs. There may be glitches here and there, depending on the state of the text you’re copying, but it will get you close, anyway.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [December 22, 2016, 9:33pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/11 "2016-12-22T21:33:21Z")

</div>

> [@dougie\_monty](#):
>
> Is that an upper-class “I” or lower-case “L”?

Look at [this](https://www.extendoffice.com/documents/word/661-replace-hard-returns-with-soft-returns.html). It’ll show you where the menus are. Just note which caret+character combos it sticks in the search box.

---

<div class="post-metadata">

**Author:** ![digs](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/digs/32/14889_2.png) [@digs](https://boards.straightdope.com/u/digs)\
**Post date:** [December 22, 2016, 10:35pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/12 "2016-12-22T22:35:58Z")

</div>

Are you pasting these into Word to print them, or read them later? If I were faced with 276 pages of text, I wouldn’t read it all despite the time I’d put in (re)formatting it.

I’m seeing if I can save you some time here. If you’re just skimming, or want to feel like you’ve done some due diligence by making an effort to read some of a proposition, why not just skim/selectively read it on the original site?

I’ve had the same formatting problems, and after spending more time formatting than reading, I was honest with myself and realized that I should just skim it instead… or search the original page for “inheritance” or whatever the issue was that I wanted to focus on.

---

<div class="post-metadata">

**Author:** ![dougie\_monty](https://avatars.discourse-cdn.com/v4/letter/d/439d5e/32.png) [@dougie\_monty](https://boards.straightdope.com/u/dougie_monty)\
**Post date:** [December 22, 2016, 11:27pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/13 "2016-12-22T23:27:34Z")

</div>

When I prepared the copy I had downloaded with Notepad, I found that it included ALL of the ballot measures from the election! :o So I deleted everything but the specific item I wanted and got it down to only six pages all told. The item in Google was somewhat misleading. Quite a saving of time and paper (and ink). 🙂

---

<div class="post-metadata">

**Author:** ![mhendo](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mhendo/32/3159_2.png) [@mhendo](https://boards.straightdope.com/u/mhendo)\
**Post date:** [December 22, 2016, 11:46pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/14 "2016-12-22T23:46:04Z")

</div>

For any of this type of stuff, on Windows, i highly recommend [Notepad++](https://notepad-plus-plus.org/).

It is so superior to plain old Windows Notepad that there’s almost no comparison.

---

<div class="post-metadata">

**Author:** ![Keeve](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@Keeve](https://boards.straightdope.com/u/Keeve)\
**Post date:** [December 23, 2016, 12:34am UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/15 "2016-12-23T00:34:40Z")

</div>

> [@dracoi](#):
>
> My guess is that you’re seeing a non-printing character for a line break. You’d have to check your MS Word view options to be sure it is displaying that character. If I’m right, you can access that character in a Find/Replace by typing ^l into Word’s search box.

Thank you so much! I’ve have the same problem as the OP for decades. I always wanted to use Find/Replace, but I could never figure out which character to put in the Find field.

> [@dracoi](#):
>
> These manual line breaks would not likely be the fault of the web programmer, but of how the copy/paste works between Word and your browser. I’ve seen this before. It’s common with PDF files as well.

I disagree. There are many times that I copy from a website and each line ends with a paragraph mark, and other times that I copy from a website and each line ends with (what I now know is) a line break. Therefore, since I am using the same browser and the same version of Word in both cases, I conclude that the difference has something to do with the website itself, and not the browser or word processor.

---

<div class="post-metadata">

**Author:** ![md2000](https://avatars.discourse-cdn.com/v4/letter/m/73ab20/32.png) [@md2000](https://boards.straightdope.com/u/md2000)\
**Post date:** [December 23, 2016, 1:28am UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/16 "2016-12-23T01:28:44Z")

</div>

If you can’t type a character, you could always copy and paste it.

I’ve done that many times with search and replace. The only downside is what happens when you replace ALL the carriage returns with space. The text may look far more dense. Maybe someone else has a better solution for preserving paragraph breaks, etc.

---

<div class="post-metadata">

**Author:** ![dougie\_monty](https://avatars.discourse-cdn.com/v4/letter/d/439d5e/32.png) [@dougie\_monty](https://boards.straightdope.com/u/dougie_monty)\
**Post date:** [December 23, 2016, 1:43am UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/17 "2016-12-23T01:43:47Z")

</div>

\<\<Are you pasting these into Word to print them, or read them later?\>\>  
I print them out; usually these are cases from California Appellate Courts or the California Supreme Court. I reformat them to save paper, ink, and space.  
The problem is that some documents seem to be formatted with each line treated as a paragraph. Hence the “broken arrow.” Now, however, I know how to overcome them… I wish I could overcome the damn ads hogging space on the right end of the screen on the SDMB. :mad:

---

<div class="post-metadata">

**Author:** ![rowrrbazzle](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/rowrrbazzle/32/414_2.png) [@rowrrbazzle](https://boards.straightdope.com/u/rowrrbazzle)\
**Post date:** [December 23, 2016, 3:16am UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/18 "2016-12-23T03:16:24Z")

</div>

> [@md2000](#):
>
> If you can’t type a character, you could always copy and paste it.
> 
> I’ve done that many times with search and replace. The only downside is what happens when you replace ALL the carriage returns with space. The text may look far more dense. Maybe someone else has a better solution for preserving paragraph breaks, etc.

It’s easy in Word. Assume all paragraphs end with x carriage returns.

1. Replace all carriage returns with some unique character string, for example "-=".

2. Then replace all x consecutive occurrences of that with one carriage return.

3. Then replace all single occurrences with a space.

If some paragraphs end with a different number of carriage returns, you may have to do some manual corrections. Or you could come up with a modification of the above steps to handle that.

---

<div class="post-metadata">

**Author:** ![Keeve](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@Keeve](https://boards.straightdope.com/u/Keeve)\
**Post date:** [December 23, 2016, 11:56am UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/19 "2016-12-23T11:56:55Z")

</div>

> [@md2000](#):
>
> If you can’t type a character, you could always copy and paste it.

One would think so, but that has not been my experience with these “broken arrows”, possibly because the “broken arrow” is not really a character, but is a visual representation of “line break”.

---

<div class="post-metadata">

**Author:** ![Keeve](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@Keeve](https://boards.straightdope.com/u/Keeve)\
**Post date:** [December 23, 2016, 12:02pm UTC](https://boards.straightdope.com/t/maddening-problem-with-downloading-text/775389/20 "2016-12-23T12:02:42Z")

</div>

If anyone wants to experiment, I found this very SD page to provide many examples. Within any individual post, there are is one paragraph mark at the very end of the post, and another after “Last edited by”; all the others are is line breaks. But there are plenty of paragraph marks elsewhere on the page.
