Linux join Command: Combine Files by a Common Field

join combines rows with the same key value from two text files into a single row. The difference from paste is that it joins based on field values, not row numbers.

What is the join command?

By default, it uses the first field of each file as the key, and both files must be sorted by that key. The default behavior is to output only matching rows. To check for missing keys as well, review the -a or -v options.

Basic syntax

join [option] file1 file2

Installed implementations and options may vary depending on the distribution. Check your system's description with man join.

Examples

Join by the first field

After sorting the two files according to the same rule, they are joined.

join names.sorted scores.sorted

Specify a delimiter

Colon-separated files must use the same delimiter system in both inputs.

join -t ':' users.sorted roles.sorted

Display unmatched rows from the first file

If you want to see unmatched rows, add -a 1.

join -a 1 names.sorted scores.sorted

Result of joining on common fields

join basically merges rows where the first field of each file is the same. The following files are pre-sorted based on the first column, which is the join key.

printf 'a Ana\nb Bob\n' > names.txt
printf 'a 7\nb 12\n' > ages.txt
join names.txt ages.txt

Example output:

a Ana 7
b Bob 12

Input files should be sorted according to the same locale and key. Rows that do not match are omitted from the default output, so to check for missing keys, review -a 1 or -a 2. The two files in the example are created in the current directory, so after the work is done, check if they are needed and clean up.

Main Options and Format

Options/Format Description
-t character Specifies the input/output field delimiter.
-1 N / -2 N Specifies the merge key fields for the first and second files.
-a 1 or -a 2 Also outputs unmatched rows from the specified file.
-v 1 or -v 2 Outputs only unmatched rows from the specified file.
-o format Specifies the output fields and order.

Precautions when using

Both files must be sorted by the same key, delimiter, and locale. Unsorted input may result in missing or incorrectly combined entries. If a key is duplicated, multiple output lines can be produced for a single key, so check the resulting count.

Frequently Asked Questions

Can an unsorted file be joined directly?

The usual join operation requires both files to be sorted by the joining field. Make a copy with the sorting order and locale matched before running.

Official Documentation

The exact behavior of the options and differences between implementations can be found in the official documentation for join.

More in This Category
Linux type Command: Identify How a Command Is Resolved

Linux type Command: Identify How a Command Is Resolved

Learn how to use the Linux type command to identify aliases, functions, built-ins, keywords, and executable paths, with essential options, practical examples, output interpretation, and common troubleshooting tips.

Linux Tutorial / File Types and Metadata: ls, file, stat, and readlink

Linux Tutorial / File Types and Metadata: ls, file, stat, and readlink

Learn what ls, file, stat, and readlink each reveal about a Linux file, including type, size, timestamps, and symbolic-link targets.

Linux wc Command: Count Lines, Words, Characters, and Bytes

Linux wc Command: Count Lines, Words, Characters, and Bytes

Learn how to use the Linux wc command to count lines, words, characters, and bytes in files or standard input, with essential options, practical examples, output interpretation, and common troubleshooting tips.

How to Check IP Addresses on Linux

How to Check IP Addresses on Linux

Learn how to check private and public IP addresses on Linux, identify the active interface and route, and distinguish IPv4, IPv6, and externally visible addresses.

Linux nslookup Command: Query DNS Names and Records

Linux nslookup Command: Query DNS Names and Records

Learn how to query DNS Names and Records with the Linux nslookup command, including practical examples, key options, and important precautions.

Linux Tutorial / PID and PPID Explained: Find and Manage Processes

Linux Tutorial / PID and PPID Explained: Find and Manage Processes

Understand what PID and PPID mean, inspect parent-child process relationships with ps and pgrep, and safely terminate a process you started yourself.

Linux ping Command: Test Host Reachability and Round-Trip Time

Linux ping Command: Test Host Reachability and Round-Trip Time

Learn how to test Host Reachability and Round-Trip Time with the Linux ping command, including practical examples, key options, and important precautions.

Linux rmdir Command: Remove Empty Directories

Linux rmdir Command: Remove Empty Directories

Learn how to remove empty directories with Linux rmdir, delete empty parent paths, diagnose failures, and understand when rm -r is different.

Linux info Command: Browse GNU Documentation

Linux info Command: Browse GNU Documentation

Learn how to use the Linux info command to browse hierarchical GNU manuals and navigate Info nodes, with essential options, practical examples, output interpretation, and common troubleshooting tips.

Linux kill Command: Send a Signal to a Process ID

Linux kill Command: Send a Signal to a Process ID

Learn how to send a Signal to a Process ID with the Linux kill command, including practical examples, key options, and important precautions.