Regular Expressions & grep
Regular Expressions
A regular expression (regex) is a compact pattern language for describing shapes of text, rather than exact literal text - "a line that starts with three digits" instead of needing to spell out every possible three-digit combination. grep is the classic tool for finding lines that match a regex, but the pattern language itself shows up across sed, awk, and many other tools too, which is why it's worth learning as its own topic.
| Pattern | Matches | |
|---|---|---|
. | Any single character | |
* | Zero or more of the previous | |
^ | Start of line | |
$ | End of line | |
[abc] | Any one of a, b, c | |
[^abc] | Any character except a, b, c | |
\{2,4\} | 2 to 4 repetitions (BRE) | |
| `\ | ` | Alternation (BRE) |
A subtlety worth sitting with: . in regex means "any character," not a literal dot - and means "zero or more of whatever came right before it," not "any characters" the way a shell wildcard does. So a.b means "an a, then any number of any characters, then a b" - a very different (and much more powerful) idea than the filename globbing you've used with ls or find.
grep flags: -i ignore case, -v invert (show non-matching lines instead), -r recursive (search every file under a directory), -n show line numbers, -c count matches instead of printing them, -E extended regex, -o print only the matched part of the line (not the whole line), -w match whole words only.
$ grep -E '^[0-9]{3}-[0-9]{4}