Text Filters
The Filter Toolbox
A filter in Unix terminology is a small program that reads text in, transforms it in one specific way, and writes text out - designed to be chained together with pipes rather than to do everything itself. This is the classic Unix philosophy: many small, sharp tools composed together, rather than one giant tool trying to do it all.
| Command | Purpose |
|---|---|
cat / tac | Print / print reversed |
head -n / tail -n | First / last N lines |
tail -f | Follow a growing file |
sort / sort -n / sort -r | Sort (numeric / reverse) |
uniq -c | Collapse & count duplicate adjacent lines |
cut -d: -f1 | Extract fields |
tr 'a-z' 'A-Z' | Translate characters |
wc -l | Count lines |
nl | Number lines |
tail -f deserves special mention: instead of printing the last lines and exiting like a normal filter, it keeps running and prints new lines as they're written - this is the standard way to watch a live log file in real time (tail -f /var/log/nginx/access.log).
$ cut -d: -f1 /etc/passwd | sort | head -3
avahi
backup
bin
Reading this pipeline left to right: cut -d: -f1 extracts just the first colon-delimited field from /etc/passwd (the username, per the account-file format from the Users & Permissions module), sort puts those usernames in alphabetical order, and head -3 keeps just the first three lines of the result - three small tools, none of which knows or cares about the others, chained into one useful answer.
Warning:uniq -conly collapses duplicate lines that are directly adjacent to each other - it has no memory of lines it saw earlier in the file. Givenapple\nbanana\napple,uniqsees three different lines because the twoapples aren't next to each other, and won't merge them. The fix, and the reason you'll see this pattern everywhere, is tosortfirst so identical lines become adjacent, then pipe intouniq -c:sort file | uniq -c.