# perl script help

**URL:** <https://boards.straightdope.com/t/perl-script-help/526662>\
**Category:** Factual Questions\
**Created:** [January 27, 2010, 5:53pm UTC](https://boards.straightdope.com/t/perl-script-help/526662 "2010-01-27T17:53:21Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 27, 2010, 5:53pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/1 "2010-01-27T17:53:21Z")

</div>

I need to process a long string that contains a bunch of keywords. Each “record” begins with a keyword, then depending on THAT keyword, other keywords follow. For example:

BEGIN PART1 blah blah blah BEGIN PART99 blah blah blah

So, how do I loop through and deal with each different sort of item? I started with a foreach loop, but then I have to keep a flag set for the type of item and that seems clumsy. Then I did a for ($i=0, $i++…) loop and that wasn’t much better. There has to be a way to deal with this, but I’m not seeing it.

Thanks!

---

<div class="post-metadata">

**Author:** ![Superfluous\_Parentheses](https://avatars.discourse-cdn.com/v4/letter/s/8edcca/32.png) [@Superfluous\_Parentheses](https://boards.straightdope.com/u/Superfluous_Parentheses)\
**Post date:** [January 27, 2010, 6:03pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/2 "2010-01-27T18:03:54Z")

</div>

I would go with something like this:

```auto

$_ = some really long string
while (length) {
  if (s/^BEGIN PART1 (\w+) (\w+) (\w+) //) {
      # do something with $1, $2 and $3
  }
  elsif (s/^BEGIN PART99 (\w+) (\w+) (\w+) //) {
      # do something with the parts
  }
  # etc
  else {
    die "No match found starting at $_";
  }
}

```

basically, this chops off any matching section of the beginning of the string in $\_ and repeats the loop until the string is empty or no match is found.

---

<div class="post-metadata">

**Author:** ![UncleRojelio](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/unclerojelio/32/3160_2.png) [@UncleRojelio](https://boards.straightdope.com/u/UncleRojelio)\
**Post date:** [January 27, 2010, 6:09pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/3 "2010-01-27T18:09:35Z")

</div>

You could use the ‘[split](http://www.comp.leeds.ac.uk/Perl/split.html)’ operator to put each delimited word into an array.

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [January 27, 2010, 6:44pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/4 "2010-01-27T18:44:57Z")

</div>

chop it up into individual records with the \*\*split \*\*operator or something similar, then pass it to an hash array with the key:value being the PARTXX:substring pair (this is assuming a 1:1 equivalence, otherwise pass to a sub-array of values) Use a **switch** conditional in a foreach loop to go through the array keys. If you have anything like a tree structure, hashes are your friend.

Any chance you can include some of the actual string, and expected outputs? Your OP’s a bit nebulous.

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 27, 2010, 6:50pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/5 "2010-01-27T18:50:30Z")

</div>

Sheesh, I gotta get more sleep. Or more caffeine. I’ll be more specific

So, I’ve already loaded everything into an array. The keyword that starts a sequence is “DEFINE” and then the type of item comes next, followed by the name of the item. I’m going to load an associative array for each item name, and depending on the type of item that array will have different data.

So:

DEFINE JUNK Junk12 stuff stuff stuff DEFINE FUBAR fubar33 foo bar etc.

JUNK type items have a different set of properties from FUBAR type items, so as soon as I find JUNK I go one way, FUBAR takes different processing. I wasn’t planning on subroutines because I’m weak on passing an array to a sub and then knowing where I was when I get back.

Maybe I create a different array for each DEFINE and then process each of those individually?

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 27, 2010, 6:52pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/6 "2010-01-27T18:52:18Z")

</div>

> [@MrDibble](#):
>
> ..Your OP’s a bit nebulous.

The original data is lengthy. I hope the above is more clear?

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 27, 2010, 7:15pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/7 "2010-01-27T19:15:48Z")

</div>

for ($i=0; i++, #line) {  
if ($line[i] =~ /^DEFINE/) {  
$i++;  
# $line[$i] is now the item type  
if ($line[$i]) =~ /JUNK/ {  
$i++;  
# now $line[$i] is the item name  
# create an assoc array for JUNK with item name as index?  
} elsif ($line[$i] =~ /FUBAR/ {  
…  
}  
# every following $line item is part of the above array, so my loop continues and I load that up  
# a new DEFINE starts a new assoc array  
}

Now loop through each assoc array?

Seems complex, but maybe it has to be.

Shoot, I don’t know how to do spacing that will show up. It looks like crap (even worse than my code generally is) left-justified.

---

<div class="post-metadata">

**Author:** ![Digital\_Stimulus](https://avatars.discourse-cdn.com/v4/letter/d/aeb1de/32.png) [@Digital\_Stimulus](https://boards.straightdope.com/u/Digital_Stimulus)\
**Post date:** [January 27, 2010, 8:46pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/8 "2010-01-27T20:46:43Z")

</div>

Just curious – why don’t you split your original data on "DEFINE "? That is:

```auto

@defines = split(/DEFINE /, $_);
foreach $define (@defines) {
   # process each DEFINE similar to your proposed if/elsifs
   # after further splitting each $define into another array
}

```

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 27, 2010, 10:25pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/9 "2010-01-27T22:25:00Z")

</div>

> [@Digital\_Stimulus](#):
>
> Just curious – why don’t you split your original data on "DEFINE "? That is:
> 
> ```auto
> 
> @defines = split(/DEFINE /, $_);
> foreach $define (@defines) {
> # process each DEFINE similar to your proposed if/elsifs
> # after further splitting each $define into another array
> }
> 
> ```

The raw data is multiple lines split in random spots. I’m current reading each line, splitting on space, and putting into one big array. I suppose I could put the array back into a string and then split on DEFINE, then split each of those on space.

Is there a way to directly split an array into multiple arrays?

---

<div class="post-metadata">

**Author:** ![Punoqllads](https://avatars.discourse-cdn.com/v4/letter/p/d2c977/32.png) [@Punoqllads](https://boards.straightdope.com/u/Punoqllads)\
**Post date:** [January 28, 2010, 12:32am UTC](https://boards.straightdope.com/t/perl-script-help/526662/10 "2010-01-28T00:32:53Z")

</div>

How about something like:

```auto

sub Dispatch( @ )
{
  my ($type, @items) = @_;

  return unless defined($type);

  if ($type eq 'JUNK')
  {
        DoJunk(@items);
  }
  elsif ($type eq 'FUBAR')
  {
        DoFubar(@items);
  }
  # etc...
  else
  {
        warn("Unrecognized type '$type'
");
  }
}

## Main starts here

my @lines = <ARGV>;

my @words = split(/\s+/, join('', @lines));

my $index;

for($index = 0; $index <= $#words; ++$index)
{
  last if $words[$index] eq 'DEFINE';
}

die "No 'DEFINE' token found in input
" if $index > $#words;

my @tokens = ();
for(; $index <= $#words; ++$index)
{
  if ($words[$index] eq "DEFINE")
  {
        Dispatch(@tokens);
        @tokens = ();
  }
  else
  {
        push(@tokens, $words[$index]);
  }
}

Dispatch(@tokens);

```

Disclaimer: no warranties expressed or implied, etc.

---

<div class="post-metadata">

**Author:** ![Digital\_Stimulus](https://avatars.discourse-cdn.com/v4/letter/d/aeb1de/32.png) [@Digital\_Stimulus](https://boards.straightdope.com/u/Digital_Stimulus)\
**Post date:** [January 28, 2010, 12:45am UTC](https://boards.straightdope.com/t/perl-script-help/526662/11 "2010-01-28T00:45:14Z")

</div>

> [@NoCoolUserName](#):
>
> The raw data is multiple lines split in random spots.

Bleahhh. By “random spots”, do you mean it’s possible to have an end of line in the middle of a word?

> [@](#):
>
> Is there a way to directly split an array into multiple arrays?

Not of which I’m aware. If there is, hopefully someone will alleviate my (our) ignorance.

---

<div class="post-metadata">

**Author:** ![Omphaloskeptic](https://avatars.discourse-cdn.com/v4/letter/o/bcef8e/32.png) [@Omphaloskeptic](https://boards.straightdope.com/u/Omphaloskeptic)\
**Post date:** [January 28, 2010, 1:33am UTC](https://boards.straightdope.com/t/perl-script-help/526662/12 "2010-01-28T01:33:57Z")

</div>

> [@NoCoolUserName](#):
>
> Is there a way to directly split an array into multiple arrays?

You can do

```auto

my @lines = <ARGV>;
my @words = map { split } @lines;

```

to avoid the string concatenation. (This assumes that a word never spans multiple lines.) Alternately, you can read the entire file into a single string, by undefining the record terminator:

```auto

{ # braces for localization of $/
  local $/;
  $_ = <ARGV>; # $_ contains the entire file (assumes a single file in @ARGV)
}
@words = split;

```

In both of these cases you are reading the entire file into memory before processing, which may be an issue if the file is very large. If memory usage is an issue you can maintain a buffer containing the last partial record read, trimming this down whenever a record becomes complete.

---

<div class="post-metadata">

**Author:** ![Punoqllads](https://avatars.discourse-cdn.com/v4/letter/p/d2c977/32.png) [@Punoqllads](https://boards.straightdope.com/u/Punoqllads)\
**Post date:** [January 28, 2010, 2:17am UTC](https://boards.straightdope.com/t/perl-script-help/526662/13 "2010-01-28T02:17:51Z")

</div>

> [@Omphaloskeptic](#):
>
> You can do
> 
> ```auto
> 
> my @lines = <ARGV>;
> my @words = map { split } @lines;
> 
> ```
> 
> to avoid the string concatenation.

Ooo, much nicer. Your ideas are intriguing and I would like to subscribe to your newsletter or service.

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 28, 2010, 4:12pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/14 "2010-01-28T16:12:17Z")

</div>

Words are not split across lines, thank bog. Sorry, “random” was a bit overenthusiastic. Records are split across lines, but at word boundaries.

> [@Omphaloskeptic](#):
>
> You can do
> 
> ```auto
> 
> my @lines = <ARGV>;
> my @words = map { split } @lines;
> 
> ```
> 
> to avoid the string concatenation. (This assumes that a word never spans multiple lines.) Alternately, you can read the entire file into a single string, by undefining the record terminator:
> 
> ```auto
> 
> { # braces for localization of $/
> local $/;
> $_ = <ARGV>; # $_ contains the entire file (assumes a single file in @ARGV)
> }
> @words = split;
> 
> ```
> 
> In both of these cases you are reading the entire file into memory before processing, which may be an issue if the file is very large. If memory usage is an issue you can maintain a buffer containing the last partial record read, trimming this down whenever a record becomes complete.

So if I give the file name as an argument, and then do \_ = &lt;ARGV&gt; I'll get the entire file in one big \_ string? Nice! Then I’ll use “split on DEFINE” and have my easy-to-process records.

Thanks!

---

<div class="post-metadata">

**Author:** ![Punoqllads](https://avatars.discourse-cdn.com/v4/letter/p/d2c977/32.png) [@Punoqllads](https://boards.straightdope.com/u/Punoqllads)\
**Post date:** [January 28, 2010, 5:39pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/15 "2010-01-28T17:39:56Z")

</div>

> [@NoCoolUserName](#):
>
> So if I give the file name as an argument, and then do \_ = &lt;ARGV&gt; I'll get the entire file in one big \_ string? Nice! Then I’ll use “split on DEFINE” and have my easy-to-process records.
> 
> Thanks!

No, $\_ = \<ARGV\> will only get the first line in the first file, or one line from standard input if no filenames are passed in on the command line.

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 28, 2010, 6:39pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/16 "2010-01-28T18:39:45Z")

</div>

> [@Punoqllads](#):
>
> No, $\_ = \<ARGV\> will only get the first line in the first file, or one line from standard input if no filenames are passed in on the command line.

Ah, but we’re )undefining the record terminator, which makes it all one, big line. (If I understand the following properly)

> [@Omphaloskeptic](#):
>
> …
> 
> ```auto
> 
> { # braces for localization of $/
> local $/;
> $_ = <ARGV>; # $_ contains the entire file (assumes a single file in @ARGV)
> }
> @words = split;
> 
> ```
> 
> …

---

<div class="post-metadata">

**Author:** ![Omphaloskeptic](https://avatars.discourse-cdn.com/v4/letter/o/bcef8e/32.png) [@Omphaloskeptic](https://boards.straightdope.com/u/Omphaloskeptic)\
**Post date:** [January 28, 2010, 6:42pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/17 "2010-01-28T18:42:41Z")

</div>

> [@Punoqllads](#):
>
> No, $\_ = \<ARGV\> will only get the first line in the first file, or one line from standard input if no filenames are passed in on the command line.

Right. But if you change the definition of the record separator / you change what Perl thinks of as the end of the line. In particular, \*\*undef /; \_=&lt;ARGV&gt;;\*\* will read to the end of the current file, not just to the first newline. The braces around this block in my code example above were so that the value of / was not changed for the rest of the file (since other places in the file may reasonably expect line-oriented reads).

(On preview, what **NoCoolUserName** said.)

---

<div class="post-metadata">

**Author:** ![Omphaloskeptic](https://avatars.discourse-cdn.com/v4/letter/o/bcef8e/32.png) [@Omphaloskeptic](https://boards.straightdope.com/u/Omphaloskeptic)\
**Post date:** [January 28, 2010, 11:16pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/18 "2010-01-28T23:16:32Z")

</div>

> [@NoCoolUserName](#):
>
> Ah, but we’re )undefining the record terminator, which makes it all one, big line. (If I understand the following properly)

One more clarification: The string $\_ will contain newlines. It won’t be “one, big line” so much as “one  
big  
string.” So you should be prepared for arguments separated with whitespace other than just spaces, when you do your record processing.

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [January 30, 2010, 4:08pm UTC](https://boards.straightdope.com/t/perl-script-help/526662/19 "2010-01-30T16:08:01Z")

</div>

> [@Omphaloskeptic](#):
>
> One more clarification: The string $\_ will contain newlines. It won’t be “one, big line” so much as “one  
> big  
> string.” So you should be prepared for arguments separated with whitespace other than just spaces, when you do your record processing.

So we add “s/  
/ /;” and life is pretty darned good.
