# perl: using split and keeping split /regex/ in returned array elements

**URL:** <https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302>\
**Category:** Factual Questions\
**Created:** [February 1, 2010, 11:23pm UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302 "2010-02-01T23:23:34Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [February 1, 2010, 11:23pm UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/1 "2010-02-01T23:23:34Z")

</div>

I want to split a string on keywords. The keywords match / [A-Z]+=/ (a space, the keyword in all caps, then an equal sign). While I want to split the string up into a nice, easy-to-use array, I need to keep the keyword with the array items. Or can I create an assoc array with the keywords as keys? That would be even better.

Thanks!

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [February 2, 2010, 12:09am UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/2 "2010-02-02T00:09:07Z")

</div>

How about something like:

```auto

#!/usr/bin/perl

use strict;
use warnings;

use Data::Dumper;

my $test = ' FOO=123 BAR=456 BAZ=789';
my $tmp = $test;
my %stuff;
while( $tmp =~ s/ ([A-Z]+)=(\w+)// ) { 
    $stuff{$1} = $2;
}

print Dumper \%stuff;

```

---

<div class="post-metadata">

**Author:** ![Punoqllads](https://avatars.discourse-cdn.com/v4/letter/p/d2c977/32.png) [@Punoqllads](https://boards.straightdope.com/u/Punoqllads)\
**Post date:** [February 2, 2010, 1:26am UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/3 "2010-02-02T01:26:02Z")

</div>

Or how about:

```auto

#!/usr/bin/perl -w

use strict;

use Data::Dumper;

my $str = " FOO=123 BAR=456 BAZ=789";
my @arr = split(/ ([A-Z]+)=/, $str);
# You have to start at index 1 because index 0 is everything
# before the first regex match
my %stuff = @arr[1 .. $#arr];

print Dumper \%stuff;

```

---

<div class="post-metadata">

**Author:** ![Pedro](https://avatars.discourse-cdn.com/v4/letter/p/919ad9/32.png) [@Pedro](https://boards.straightdope.com/u/Pedro)\
**Post date:** [February 2, 2010, 2:29am UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/4 "2010-02-02T02:29:15Z")

</div>

I just started learning perl so I’ll take another shot at it, in case you want positions too (otherwise a hash is definitely the way to go).

```auto

#!/usr/bin/perl -w

use strict;

use Data::Dumper;

my $str = " FOO=123 BAR=456 BAZ=789";
my @stuff;

foreach ( split(/\s/, $str) ) {
    if ( /([A-Z]+)=(\w+)/ ) {
        push @stuff, ([$1, $2]);
    }
}

print Dumper (\@stuff);

```

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [February 2, 2010, 3:51pm UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/5 "2010-02-02T15:51:34Z")

</div>

> [@Punoqllads](#):
>
> Or how about:
> 
> ```auto
> 
> #!/usr/bin/perl -w
> 
> use strict;
> 
> use Data::Dumper;
> 
> my $str = " FOO=123 BAR=456 BAZ=789";
> my @arr = split(/ ([A-Z]+)=/, $str);
> # You have to start at index 1 because index 0 is everything
> # before the first regex match
> my %stuff = @arr[1 .. $#arr];
> 
> print Dumper \%stuff;
> 
> ```

You saved the keyword using ([A-Z]+), but how did it get into %stuff? I mean, obviously @arr[1 .. $#arr] does it, but what happens there?

Zowie, that’s pretty spiff.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [February 2, 2010, 4:12pm UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/6 "2010-02-02T16:12:39Z")

</div>

split has a little known feature that causes capture-groups in the regex to also be returned along with the split pieces. From the [perldoc](http://perldoc.perl.org/functions/split.html):

> [@](#):
>
> If the PATTERN contains parentheses, additional list elements are created from each matching substring in the delimiter.

So @arr will contain ( ‘’, ‘FOO’, ‘123’, ‘BAR’, ‘456’, ‘BAZ’, ‘789’ )

Whenever you have alternating keys and values in a list, you can just assign that to a hash. Though it’s a little prettier to say

```auto

shift @arr;
my %stuff = @arr;

```

---

<div class="post-metadata">

**Author:** ![NoCoolUserName](https://avatars.discourse-cdn.com/v4/letter/n/5fc32e/32.png) [@NoCoolUserName](https://boards.straightdope.com/u/NoCoolUserName)\
**Post date:** [February 2, 2010, 6:18pm UTC](https://boards.straightdope.com/t/perl-using-split-and-keeping-split-regex-in-returned-array-elements/527302/7 "2010-02-02T18:18:43Z")

</div>

> [@friedo](#):
>
> split has a little known feature that causes capture-groups in the regex to also be returned along with the split pieces…

Woo-hoo! That’s what I was looking for!
