Linux Tutorial / Search and Process Text with grep, sort, uniq, cut, and wc
Use grep to select matching lines, sort to order lines, uniq to combine adjacent duplicates, cut to extract delimited fields, and wc to count lines, words, or bytes. These commands usually write a result to standard output rather than editing the input file. By connecting them in the right order, you can perform useful small text-processing tasks.
Create a small data file
This lesson assumes Bash and GNU utilities. If ~/linux-course-12-practice already exists, choose another name. Do not continue with file creation if the new directory cannot be made: > overwrites an existing target file.
mkdir ~/linux-course-12-practice cd ~/linux-course-12-practice pwd
After confirming your location, create five records with a name and status separated by a colon.
printf 'alpha:ready\nbeta:waiting\nalpha:done\ngamma:ready\nbeta:ready\n' > records.txt cat records.txt
cat prints the five lines you entered. Every example below uses this same file.
Select matching lines with grep
grep prints lines that match a pattern. Add -n to show each line's number in the input file.
grep -n 'ready' records.txt
Three lines match; waiting does not contain ready.
1:alpha:ready 4:gamma:ready 5:beta:ready
A basic grep pattern is a regular expression, so check the meaning of special characters. Use grep -F for a literal fixed string. No matching line generally gives exit status 1, while a command error gives 2; do not treat all empty output as the same failure.
Extract fields with cut and count with wc
cut -d: -f1 takes the first colon-delimited field from each record, leaving only names:
cut -d: -f1 records.txt
alpha beta alpha gamma beta
cut handles simple delimiters; it does not parse quoted fields or embedded line breaks as a full CSV parser would. wc -l records.txt shows a line count of 5 beside the filename.
wc -l records.txt
wc -w counts words, and wc -c counts bytes. A multibyte character such as a Korean character can occupy more than one byte. GNU wc -m counts characters when that is what you need.
Connect sort and uniq in the right order
sort orders lines. uniq combines only adjacent identical lines, so sort first when identical names are separated in the input. The pipe | passes one command's standard output to the next command's standard input.
cut -d: -f1 records.txt | sort | uniq -c
The result counts alpha twice, beta twice, and gamma once. Leading spaces and sort order can vary with locale and implementation.
2 alpha
2 beta
1 gamma
If you run sort | uniq on the full records instead, alpha:ready and alpha:done remain distinct. Decide what constitutes a duplicate before choosing which field to extract. Sorting uses locale collation rules unless you specify otherwise, so compare locale settings when results differ between systems.
Choose a tool for the question
| Goal | Tool | Key consideration |
|---|---|---|
| Matching lines | grep |
Regular expression versus literal pattern |
| Line ordering | sort |
Locale and sort key |
| Deduplication or frequency | uniq |
Equal values must be adjacent |
| Simple delimited fields | cut |
Delimiter and field number |
| Lines, words, or bytes | wc |
Choose the counting option |
If a result is empty or unexpected, first inspect records.txt with cat. Then run each part of a longer pipeline separately to see whether the pattern, delimiter, or sorting stage changed the result.
Official documentation
The GNU grep manual covers matching and exit statuses; the GNU Coreutils manual covers sort, uniq, cut, and wc.









