# Refresh my memory on fscanf (C programming)

**URL:** <https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489>\
**Category:** Factual Questions\
**Created:** [January 9, 2019, 9:44pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489 "2019-01-09T21:44:52Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [January 9, 2019, 9:44pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/1 "2019-01-09T21:44:52Z")

</div>

I’m sorry to ask such a trivial question, but it’s been a long time since I’ve done any file-input programming, and a longer time since my programming classes.

I have a program I’m writing in C, which is reading input from a file. I want to read in everything in the file (across multiple lines, if it matters) until I get to a particular string, and discard all of that. Then, only after I’ve read in that particular string, I’ll start actually reading in the real data that I want. I thought I could do something like

```auto

fscanf(infile,"STARTHERE");

```

but that only moves the stream so far as the start of the file matches the string.

Then I tried

```auto

fscanf(infile,"%sSTARTHERE",junk);

```

(where junk is, of course, a previously-defined string variable), but that just puts until the first whitespace into the junk string, and does nothing with the STARTHERE string.

I’m sure there’s a simple way to do this, but I’m blanking on what it is.

(note: The string I’m actually searching for isn’t literally “STARTHERE”; I’m just using that in the question for simplicity)

---

<div class="post-metadata">

**Author:** ![beowulff](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/beowulff/32/542_2.png) [@beowulff](https://boards.straightdope.com/u/beowulff)\
**Post date:** [January 9, 2019, 10:31pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/2 "2019-01-09T22:31:04Z")

</div>

I don’t think you can do this in one pass with fscanf.

Do it in two passes - scan until you get to your match, then scan again, with your %s conversion.

---

<div class="post-metadata">

**Author:** ![Marvin\_the\_Martian](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/marvin_the_martian/32/2898_2.png) [@Marvin\_the\_Martian](https://boards.straightdope.com/u/Marvin_the_Martian)\
**Post date:** [January 9, 2019, 10:34pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/3 "2019-01-09T22:34:41Z")

</div>

[This might help.](https://www.codingunit.com/c-tutorial-searching-for-strings-in-a-text-file)

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 9, 2019, 10:35pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/4 "2019-01-09T22:35:19Z")

</div>

Here is one possibility: [c - strstr on huge mmapped file - Stack Overflow](https://stackoverflow.com/questions/41187170/strstr-on-huge-mmapped-file)

> [@beowulff](#):
>
> I don’t think you can do this in one pass with fscanf.
> 
> Do it in two passes - scan until you get to your match, then scan again, with your %s conversion.

Or, don’t use fscanf

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 9, 2019, 10:42pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/5 "2019-01-09T22:42:09Z")

</div>

> [@Marvin\_the\_Martian](#):
>
> [This might help.](https://www.codingunit.com/c-tutorial-searching-for-strings-in-a-text-file)

Looks obviously buggy since their algorithm is to read 512 bytes at a time and search each segment; what if the string falls across a boundary?

---

<div class="post-metadata">

**Author:** ![Jas09](https://avatars.discourse-cdn.com/v4/letter/j/d07c76/32.png) [@Jas09](https://boards.straightdope.com/u/Jas09)\
**Post date:** [January 9, 2019, 10:44pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/6 "2019-01-09T22:44:35Z")

</div>

That’s correct (that you can’t do it with one call to fscanf). In fact, if STARTHERE could be embedded within a larger string (i.e. not delimited by whitespace) then it’s a bit more difficult that it would seem.

Is it correct that the text prior to STARTHERE could contain any combination of any-length strings plus white space (including new lines)?

If so, I think you will need to call fscanf(infile,"%s",junk) in a while loop and then do a strstr(junk, “STARTHERE”) call to see if the string you read in contains your starting string. If yes, exit the while loop. Unfortunately, some of the data you want is now embedded in your junk variable, so you’ll need to parse that part (strstr returns a pointer to the matching string) before reading in the rest of the data.

On preview: **Marvin** ’s link has the code for what I’m trying to describe, but uses fgets() instead of fscanf(), which is possibly better. But as **DPRK** points out, if you can’t assume that your lines are \<512 bytes or that your token can’t cross that boundary then it won’t work. The most reliable way is actually probably to either read the whole file or do it one character at a time until you find the well-formed part of your file.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [January 10, 2019, 2:17am UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/7 "2019-01-10T02:17:45Z")

</div>

> [@](#):
>
> Quoth **beowulf** :
> 
> I don’t think you can do this in one pass with fscanf.
> 
> Do it in two passes - scan until you get to your match, then scan again, with your %s conversion.

Yeah, that’s what I’m trying to do (except the actual data will mostly be %f conversion, not %s), but I’m not sure how to do the “scan until I get to my match” part.

As of right this moment, the full text of my input file up to the start point is probably consistent, but I don’t want to count on it remaining so indefinitely (there’s a lot of header-type data, which includes a couple of version numbers which will eventually change). On the other hand, my STARTHERE text will _probably_ always have whitespace preceding it and ending it, and will always have whitespace within it (so I suppose I could just treat it as two strings, since the part up to the first whitespace should be distinctive enough).

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 10, 2019, 2:39am UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/8 "2019-01-10T02:39:01Z")

</div>

You should be able to sscanf a memory-mapped file, after searching for your STARTHERE, in order to read your %f formatted data.

---

<div class="post-metadata">

**Author:** ![beowulff](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/beowulff/32/542_2.png) [@beowulff](https://boards.straightdope.com/u/beowulff)\
**Post date:** [January 10, 2019, 2:39am UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/9 "2019-01-10T02:39:43Z")

</div>

If the header data is variable, you probably should use a regex search.  
Or, read your input into a buffer, then use strstr() to find the ‘STARTHERE’ string, then parse the rest of string with scanf().

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 10, 2019, 2:48am UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/10 "2019-01-10T02:48:59Z")

</div>

strstr does not support regular expressions, but naturally there are ready libraries that do, if needed.

---

<div class="post-metadata">

**Author:** ![Francis\_Vaughan](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/francis_vaughan/32/3093_2.png) [@Francis\_Vaughan](https://boards.straightdope.com/u/Francis_Vaughan)\
**Post date:** [January 10, 2019, 4:33am UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/11 "2019-01-10T04:33:07Z")

</div>

> [@DPRK](#):
>
> You should be able to sscanf a memory-mapped file, after searching for your STARTHERE, in order to read your %f formatted data.

This.  
There no reason not to memory map the file unless it is a stream. You pay a silly amount of overhead for no good reason by not mapping it. It isn’t even as if this is a new thing. I was doing this in about 1985. And if you need a regex, the world is your oyster.

---

<div class="post-metadata">

**Author:** ![Quercus](https://avatars.discourse-cdn.com/v4/letter/q/7ab992/32.png) [@Quercus](https://boards.straightdope.com/u/Quercus)\
**Post date:** [January 10, 2019, 3:34pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/12 "2019-01-10T15:34:25Z")

</div>

Alternatively, couldn’t you just read the file one character at a time checking for a match?

I’m rusty on C so treat this as pseudo-code, but something like: [Grrr, sorry about the lack of indentation]

Found = FALSE  
C=1 /\* character in the Target$ that we’re checking for a match  
Target$=\<what you’re looking for\>

Do While Not Found  
B$ = \<read next character from file\>  
if B$ \<\>Target$[C] then  
C=1  
else  
C=C+1  
If C \> Length (Target$) then  
Found =TRUE  
endif  
endif  
loop

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 10, 2019, 6:43pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/13 "2019-01-10T18:43:20Z")

</div>

> [@Quercus](#):
>
> Alternatively, couldn’t you just read the file one character at a time checking for a match?
> 
> I’m rusty on C so treat this as pseudo-code, but something like: [Grrr, sorry about the lack of indentation]
> 
> Found = FALSE  
> C=1 /\* character in the Target$ that we’re checking for a match  
> Target$=\<what you’re looking for\>
> 
> Do While Not Found  
> B$ = \<read next character from file\>  
> if B$ \<\>Target$[C] then  
> C=1  
> else  
> C=C+1  
> If C \> Length (Target$) then  
> Found =TRUE  
> endif  
> endif  
> loop

Don’t trust the standard library implementation? 🙂

You can do better than that, though, like with a Boyer–Moore algorithm.

---

<div class="post-metadata">

**Author:** ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)\
**Post date:** [January 10, 2019, 7:07pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/14 "2019-01-10T19:07:52Z")

</div>

(also, the pseudo code above is buggy)

---

<div class="post-metadata">

**Author:** ![rat\_avatar](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/rat_avatar/32/255_2.png) [@rat\_avatar](https://boards.straightdope.com/u/rat_avatar)\
**Post date:** [January 10, 2019, 11:13pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/15 "2019-01-10T23:13:02Z")

</div>

The Variadic Functions like the scanf family in the standard library are unsafe and should never be trusted for input you don’t control. That said a while loop with sentinel values is the typical solution even using them.

That said, you use ‘\*’ to suppress assignment of conversions, so:

scanf("%c %\*c",&a);

Will discard the second char in that matched conversion, or you can match in a while/for loop.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [January 10, 2019, 11:46pm UTC](https://boards.straightdope.com/t/refresh-my-memory-on-fscanf-c-programming/827489/16 "2019-01-10T23:46:59Z")

</div>

OK, thanks, everyone, I’ve got it working. The method I ended up going with was

```auto

  searching = 1;
  while(searching) {
    fscanf(infile,"%s",junk);
    if(strstr("START",junk)) searching = 0;
  }
  fscanf(infile," HERE");

```

(where “START HERE” is my two-word keystring).

> [@](#):
>
> Quoth **Francis Vaughan** :
> 
> There no reason not to memory map the file unless it is a stream. You pay a silly amount of overhead for no good reason by not mapping it. It isn’t even as if this is a new thing. I was doing this in about 1985.

Well, there is another reason: I never learned to do it before (back in the mid 90s, while it might have been possible, stupid memory limits often made it difficult), and the extra computer time consumed by the silly amount of overhead, even over this program’s entire expected usage life, is less than the extra programmer time it’d take me to learn it that way. And it’s clearly not something that comes up in very many of my other programs, given how out of practice I was on file IO to begin with.

Now, there _are_ some areas of this code (after the IO part I have to get through, before I even have a chance to play around with the fun parts) where I am in fact worried about efficiency, but those are parts that involve (admittedly low precision) trig functions, and which might end up getting looped over a million times, unless I can think of some more clever way than brute force for those algorithms.
