# Simple PERL programming question

**URL:** <https://boards.straightdope.com/t/simple-perl-programming-question/355287>\
**Category:** Factual Questions\
**Created:** [May 3, 2006, 8:11pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287 "2006-05-03T20:11:50Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![pulykamell](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pulykamell/32/3166_2.png) [@pulykamell](https://boards.straightdope.com/u/pulykamell)\
**Post date:** [May 3, 2006, 8:11pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/1 "2006-05-03T20:11:50Z")

</div>

This is probably trial, but I’m pretty new to PERL, and I can’t figure out what the neatest way to accomplish the following is:

I have one file that is a list of strings that need to be found. Let’s call it KEYWORDS.TXT. The file contains about 100 or so keywords that are listed in the format KEYWORD  
. (That is, simply a word followed by a newline character).

I have another file that is a list that the text that needs to be searched. Let’s call it TEXT.TXT.

What I need to do is iterate over each line of text, and check whether it contains any of the keywords. So, basically, while (\<INFILE\>) if $\_ contains an item from KEYWORDS.TXT print OUTFILE, else loop.

So, do I need to read KEYWORDS.TXT into an array and iterate through each item of the array for every cycle of the main loop? I’m not quite sure how all the I/O works in PERL, so I don’t know how to read this into an array.

Or is there some simpler solution I’m missing. I’m fairly new to this, so go easy on me. 🙂

---

<div class="post-metadata">

**Author:** ![dre2xl](https://avatars.discourse-cdn.com/v4/letter/d/34f0e0/32.png) [@dre2xl](https://boards.straightdope.com/u/dre2xl)\
**Post date:** [May 3, 2006, 8:18pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/2 "2006-05-03T20:18:22Z")

</div>

Well, you could do the double iteration, but another way is to read KEYWORDS.TXT into a hash.

Such as this:

my $keywordsHash;  
$keywordsHash{$keywordsFileLines[1]} = 1;  
$keywordsHash{$keywordsFileLines[2]} = 1;

and so on.

Then, when you’re reading TEXT.TXT, do this:

foreach (@textFile)  
{  
if (keywordsHash{\_})  
{  
[print OUTFILE]  
}  
}

---

<div class="post-metadata">

**Author:** ![Meros](https://avatars.discourse-cdn.com/v4/letter/m/8797f3/32.png) [@Meros](https://boards.straightdope.com/u/Meros)\
**Post date:** [May 3, 2006, 8:59pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/3 "2006-05-03T20:59:44Z")

</div>

> [@dre2xl](#):
>
> Well, you could do the double iteration, but another way is to read KEYWORDS.TXT into a hash.
> 
> \<snip…\>
> 
> foreach (@textFile)  
> {  
> if (keywordsHash{\_})  
> {  
> [print OUTFILE]  
> }  
> }

It looks like that would only work if your text.txt had one word per line. Otherwise, the key for your hash would contain the entire line and you would end up with a lot of false results.

If there’s more than one word per line of the file to be scanned, then I would do a double iteration. (please ignore my sloppy code, i’m writing this quickly, god i love perl 😃 )

```auto

open FH, "<keywords.txt";
$i=0;

while(<FH>){
    $keywords[$i]=chomp($_);
    $i++;
}
close FH;

```

Do that to load your keywords into an array to step through, then for the other file, for the purposes of this i’m going to dump lines which match into another variable for writing to a file

```auto

open FH, "<text.txt";
while(<FH>){

    $line=$_;
     foreach $key (@keywords){
         $pat="/$key/";
         
         if($line =~ $pat){
               $newtxt=$newtxt.$line;
         }
    }
}

```

Again, this as it stands may or may not be functional (haven’t tested it, but i’ve been writing code like this all day), but hopefully I got the gist across.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [May 3, 2006, 9:05pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/4 "2006-05-03T21:05:25Z")

</div>

> [@pulykamell](#):
>
> What I need to do is iterate over each line of text, and check whether it contains any of the keywords. So, basically, while (\<INFILE\>) if $\_ contains an item from KEYWORDS.TXT print OUTFILE, else loop.

First iterate over the keywords and build an array containing each keyword as an element:

```auto

use strict;
use warnings;

my @keywords;
open my $kwfh, "KEYWORDS.TXT" or die $!;
while( <$kwfh> ) { 
    chomp;
    push @keywords, $_;
}

```

Then iterate over the lines in TEXT and check to see if each keyword appears:

```auto

open my $tfh, "TEXT.TXT" or die $!;
while( <$tfh> ) { 
    foreach my $kw( @keywords ) { 
        if( /\Q$kw/ ) { 
            print;
            last;
        }
    }
}

```

---

<div class="post-metadata">

**Author:** ![dre2xl](https://avatars.discourse-cdn.com/v4/letter/d/34f0e0/32.png) [@dre2xl](https://boards.straightdope.com/u/dre2xl)\
**Post date:** [May 3, 2006, 9:10pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/5 "2006-05-03T21:10:34Z")

</div>

Yeah, I was picturing one word per line in TEXT.txt . for multiple words per line, your solution’s good =)

---

<div class="post-metadata">

**Author:** ![pulykamell](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pulykamell/32/3166_2.png) [@pulykamell](https://boards.straightdope.com/u/pulykamell)\
**Post date:** [May 3, 2006, 9:21pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/6 "2006-05-03T21:21:35Z")

</div>

> [@friedo](#):
>
> First iterate over the keywords and build an array containing each keyword as an element:
> 
> ```auto
> 
> use strict;
> use warnings;
> 
> my @keywords;
> open my $kwfh, "KEYWORDS.TXT" or die $!;
> while( <$kwfh> ) { 
> chomp;
> push @keywords, $_;
> }
> 
> ```
> 
> Then iterate over the lines in TEXT and check to see if each keyword appears:
> 
> ```auto
> 
> open my $tfh, "TEXT.TXT" or die $!;
> while( <$tfh> ) { 
> foreach my $kw( @keywords ) { 
> if( /\Q$kw/ ) { 
> print;
> last;
> }
> }
> }
> 
> ```

I tried it once myself, and I tried it once using your exact code.

For some reason, for both your & my program, it’s spitting everything out at me. Every single line matches. I can’t for the life of me figure out why.

I commented out the second half of the program, to make sure the array is being read in right, and the array itself is fine. 134 items, each one consisting of one element, everything is okay.

So something in the for loop is causing everything to match.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [May 3, 2006, 9:31pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/7 "2006-05-03T21:31:09Z")

</div>

Does your KEYWORDS.TXT have any blank lines? If so, an empty-string as a keyword would match everything. You could modify the loop to read the keywords to filter them out:

```auto

my @keywords;
open my $kwfh, "KEYWORDS.TXT" or die $!;
while( <$kwfh> ) { 
    chomp;
    next unless length $_;
    push @keywords, $_;
}

```

---

<div class="post-metadata">

**Author:** ![pulykamell](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pulykamell/32/3166_2.png) [@pulykamell](https://boards.straightdope.com/u/pulykamell)\
**Post date:** [May 3, 2006, 9:36pm UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/8 "2006-05-03T21:36:38Z")

</div>

> [@friedo](#):
>
> Does your KEYWORDS.TXT have any blank lines? If so, an empty-string as a keyword would match everything. You could modify the loop to read the keywords to filter them out:
> 
> ```auto
> 
> my @keywords;
> open my $kwfh, "KEYWORDS.TXT" or die $!;
> while( <$kwfh> ) { 
> chomp;
> next unless length $_;
> push @keywords, $_;
> }
> 
> ```

:smack:

I had a blank line at the very end of the file. See, this is why I’m not a programmer. Stuff like this would drive me bonkers. Last time I was in this position, it was substituting a “=” for a “==”. Took me an hour to figure out. Only to figure out that I was _still_ wrong and needed an “eq”.

Argh.

Thanks all!

---

<div class="post-metadata">

**Author:** ![Punoqllads](https://avatars.discourse-cdn.com/v4/letter/p/d2c977/32.png) [@Punoqllads](https://boards.straightdope.com/u/Punoqllads)\
**Post date:** [May 4, 2006, 12:46am UTC](https://boards.straightdope.com/t/simple-perl-programming-question/355287/9 "2006-05-04T00:46:41Z")

</div>

You can do it without a nested loop using slices if the files are small enough to fit into memory.

```auto

open(KEYWORDS, "KEYWORDS.TXT");
open(TEXTFILE, "TEXT.TXT");
open(OUTFILE, ">OUTPUT.TXT");

my (@keywords, @lines, %lineNum, %outlines);

@keywords = <KEYWORDS>;
@lines = <TEXTFILE>;

# Note that it's "@lineNum", not "%lineNum" or even "$lineNum".
# We're using slices.

# lineNum is so that we can print out the valid lines in the same order they
# were in the input text file.
@lineNum{@lines} = (0 .. $#lines);

chomp @keywords;

foreach $word (@keywords)
{
  # See previous note on slices.
  @outlines{grep(/$word/, @lines)} = 1;
}

print( OUTFILE sort { $lineNum{$a} <=> $lineNum{$b} } keys %outlines );

```
