Linux split command

Linux 命令大全Linux Command Reference

splitsplit is a file splitting tool built into Linux that can split a file by line count, size, or a specified quantity.

splitThe command is used to split a large file into several smaller files, making it easy to transfer, store, or process in parallel.

The split files can be reassembled using thecatcommand to merge them back into the original file.

When log files are too large to open with an editor, when a large file needs to be uploaded in chunks, or when data needs to be processed in parallel,splitsplit is the preferred solution.

Split files are named by default withxaa、xab、xacand other alphabetic sequences.

splitsplit does not delete or modify the original file, so the split operation is safe. However, when splitting into many small files, make sure there is sufficient disk space.

Command syntax

splitThe basic syntax of split is as follows:

split [选项] [输入文件] [输出文件前缀]

If no input file is specified, data is read from standard input.

If no output file prefix is specified, the default prefix isxused.

splitThe common options of split are listed below:

OptionDescriptionExample
-b, --bytes=大小Split the file by the specified number of bytessplit -b 100M large.log
-l, --lines=行数Split the file by the specified number of linessplit -l 1000 data.txt
-n, --number=数量Split into the specified number of small filessplit -n 5 data.txt
-a, --suffix-length=NSpecify the suffix length (default is 2)split -a 3 -l 100 data.txt
-d, --numeric-suffixesUse numeric suffixes instead of alphabetic suffixessplit -d -l 100 data.txt
-C, --line-bytes=大小Split by size while keeping each line completesplit -C 10M log.txt
--verboseDisplay detailed information about the split processsplit --verbose -b 10M file.bin

Supported size units include:K(KB)、M(MB)、G(GB)、T(TB) and other suffixes.

Difference between -b and -C:

OptionBehaviorApplicable scenario
-bStrictly cuts by byte count and may truncate in the middle of a lineBinary files or scenarios where line breaks do not matter
-CTries to cut at line boundaries when reaching the size limit, ensuring each line is completeText logs, CSV, and other scenarios where lines must remain intact

When splitting text files, prefer-Cover-bto avoid a line being cut off and split between two files.


Detailed usage

Split by line count

Splitting by line count is the most intuitive way and is suitable for processing log files or CSV data.

# 生成一个包含 5000 行的测试文件
$ seq 1 5000 > data.txt

# 每 1000 行拆分为一个小文件
$ split -l 1000 data.txt part_

# 查看生成的文件
$ ls -lh part_*
-rw-r--r-- 1 example example 3.9K May 19 14:30 part_aa
-rw-r--r-- 1 example example 3.9K May 19 14:30 part_ab
-rw-r--r-- 1 example example 3.9K May 19 14:30 part_ac
-rw-r--r-- 1 example example 3.9K May 19 14:30 part_ad
-rw-r--r-- 1 example example 3.9K May 19 14:30 part_ae

After running, 5 files are generated, each with exactly 1000 lines, and the filenames usepart_as the prefix.

Split by file size

Splitting by size is suitable for cutting binary files or limiting the maximum size of a single file.

# 生成一个 10MB 的测试文件
$ dd if=/dev/urandom of=test.bin bs=1M count=10

# 按每 2MB 拆分
$ split -b 2M test.bin chunk_

# 查看拆分结果
$ ls -lh chunk_*
-rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_aa
-rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ab
-rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ac
-rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ad
-rw-r--r-- 1 example example 2.0M May 19 14:32 chunk_ae

Specify the number of splits

Use-nto divide the file evenly into the specified number of small files.

# 将文件均分为 3 份
$ split -n 3 data.txt equal_

# 查看各文件行数
$ wc -l equal_*
  1667 equal_aa
  1667 equal_ab
  1666 equal_ac
  5000 total

The file is distributed as evenly as possible; a file with a total of 5000 lines is divided into about 1667 lines per file.

When the total number of lines is not evenly divisible, the earlier files get a few extra lines and the later files get a few fewer.

Use numeric suffixes

The default alphabetic suffix (aa、ab...) is not intuitive enough; you can use-dto switch to numeric suffixes.

# 使用数字后缀,并指定后缀长度为 3
$ split -d -a 3 -l 1000 data.txt example_

# 生成的文件名
$ ls example_*
example_000  example_001  example_002  example_003  example_004

Here, a suffix length of 3 means up to 1000 files are supported (from 000 to 999). When splitting into many small files, you can increase the value of-athe option.

Split text while preserving line integrity

Use-Cto split by size while ensuring that no line is truncated.

# 生成一个包含不同长度行的测试日志
$ for i in $(seq 1 100); do echo "Line $i: EXAMPLE testing data $(head -c $((RANDOM % 50 + 10)) /dev/urandom | base64)"; done > log.txt

# 按 1KB 拆分,同时保持行完整性
$ split -C 1K log.txt log_

# 检查每个文件的最后一行是否完整(都以换行符结尾)
$ for f in log_*; do echo "$f: $(tail -c 1 $f | xxd | grep -c '0a')"; done

-CThis ensures that the last line of each split file is complete, and lines are never cut in half.

Merge split files

The split files can be merged losslessly to restore the original file:

# 合并所有拆分文件,还原为原始文件
$ cat part_* > restored.txt

# 验证合并后的内容与原始文件一致
$ diff data.txt restored.txt && echo "文件一致,合并成功"
文件一致,合并成功

When merging,catcat concatenates them in the alphabetical order of shell wildcard expansion, which is exactly thesplitorder in which split generates the files.


FAQ

The split files still occupy the same total disk space as the original file (actually, there is some additional filesystem metadata overhead), so make sure there is enough free disk space.

When using numeric suffixes, if the number of split files exceeds the maximum representable by the suffix length (e.g.,-a 2up to 100 files),splitsplit will exit with an error. Increase the value of the-aparameter to fix it.

If you need to split by specific content (such as by a delimiter),splitsplit does not support it; use thecsplitcsplit command instead.

When reading data from standard input,splitsplit cannot estimate the total size, so the-noption for specifying the number of splits does not support standard input mode.


Related commands

CommandDescription
catMerge files or output file contents
csplitSplit files by content (regular expression)
wcCount the lines, words, and bytes in a file
ddConvert and copy files, and extract data by block size
headOutput the beginning of a file
tailOutput the end of a file

Linux 命令大全Linux Command Reference

Other extensions