Awk has several very powerful built-in variables. Generally speaking, they are divided into two types: - The first type are variables that can be changed, such as field separator (FS) and record separator (RS) - The second type can be used for data processing or data summarization, such as record number (NR) and field number (NF). This article introduces:FS,OFS, RS, ORS, NR, NR, FNR
FS: Input Field Separator Variable
FS(Field Separator) When reading and parsing each line in the input file, by default it is split by spaces into field variables, $1, $2, etc.FSThe variable is used to set the field separator symbol for each record.FSIt can be any string or regular expression. You can declare FS in the following two ways:
Use-Fcommand-line option
using it as an ordinary variable assignment
语法:
$ awk -F 'FS' 'commands' inputfilename
或者
$ awk 'BEGIN{FS="FS";}'
FSIt can be any character or regular expression.
FSIt can be changed multiple times, but remains unchanged until explicitly modified. However, if you want to change the field separator, it is best to change it before reading the text,FSso that the change will take effect on the text you read in.
Below is an example of usingFSto read /etc/passwd with:as the separator
$ cat etc_passwd.awk
BEGIN{
FS=":";
print "Name\tUserID\tGroupID\tHomeDirectory";
}
{
print $1"\t"$3"\t"$4"\t"$6;
}
END {
print NR,"Records Processed";
}
Output:
$ awk -f etc_passwd.awk /etc/passwd Name UserID GroupID HomeDirectory gnats 41 41 /var/lib/gnats libuuid 100 101 /var/lib/libuuid syslog 101 102 /home/syslog hplip 103 7 /var/run/hplip avahi 105 111 /var/run/avahi-daemon saned 110 116 /home/saned pulse 111 117 /var/run/pulse gdm 112 119 /var/lib/gdm 8 Records Processed
OFS: Output Field Separator Variable
OFS(Output Field Separator) is equivalent toFSFS on output. By default, a space character is used as the output separator. Below is anOFSexample:
$ awk -F':' '{print $3,$4;}' /etc/passwd
41 41
100 101
101 102
103 7
105 111
110 116
111 117
112 119
Note that the comma in the print statement in the command means using a space to connect two parameters, which is the default OFS value. ThereforeOFSit can be inserted between output fields as shown below:
$ awk -F':' 'BEGIN{OFS="=";} {print $3,$4;}' /etc/passwd
41=41
100=101
101=102
103=7
105=111
110=116
111=117
112=11
RS: Record Separator
RS(Record Separator) defines a record line. When reading a file, by default one line is treated as one record. The following example uses student.txt as the input file, with records separated by two blank lines, and each field of each record separated by a newline:
$ cat student.txt Jones 2143 78 84 77 Gondrol 2321 56 58 45 RinRao 2122 38 37 65 Edwin 2537 78 67 45 Dayan 2415 30 47 20
Then the following script will output two records from student.txt:
$ cat student.awk
BEGIN {
RS="\n\n";
FS="\n";
}
{
print $1,$2;
}
$ awk -f student.awk student.txt
Jones 2143
Gondrol 2321
RinRao 2122
Edwin 2537
Dayan 2415
In student.awk, each student's detailed information is treated as one record, because RS (record separator) is set to two newlines. And becauseFS(field separator) is a newline, so one line is one field.
ORS: Output Record Separator Variable
ORS(Output Record Separator) as the name implies, is equivalent toRSRS for output. Each record will be separated by a separator when output. See the followingORSexample:
$ awk 'BEGIN{ORS="=";} {print;}' student-marks
Jones 2143 78 84 77=Gondrol 2321 56 58 45=RinRao 2122 38 37 65=Edwin 2537 78 67 45=Dayan 2415 30 47 20=
In the above script, each record of the input file is=separated. Note: student-marks is the output above.
NR: Number of Records Variable
NR(Number of Record) indicates the total number of records already processed, or the line number (not necessarily of one file; it may be multiple). In the example below,NRit represents the line number; in the END part,NRit is the total number of all records in the files.
$ awk '{print "Processing Record - ",NR;}END {print NR, "Students Records are processed";}' student-marks
Processing Record - 1
Processing Record - 2
Processing Record - 3
Processing Record - 4
Processing Record - 5
5 Students Records are processed
NF: Number of Fields in a Record
NF(Number for Field) indicates the number of fields in a record. It is very useful in determining whether all fields exist in a certain record. Let's look at the student-mark file as follows:
$ cat student-marks Jones 2143 78 84 77 Gondrol 2321 56 58 45 RinRao 2122 38 37 Edwin 2537 78 67 45 Dayan 2415 30 47
Then the following Awk program prints the record number (NR), and the number of fields in that record: Therefore it is very easy to discover which data is missing.
$ awk '{print NR,"->",NF}' student-marks
1 -> 5
2 -> 5
3 -> 4
4 -> 5
5 -> 4
FILENAME: Name of the Current Input File
FILENAMERepresents the name of the file currently being read. AWK can accept and read many files for processing. See the following example:
$ awk '{print FILENAME}' student-marks
student-marks
student-marks
student-marks
student-marks
student-marks
The name is output for every record in the input file.
FNR: Number of Records in the Current Input File
When awk reads multiple files,NRit represents the total number of records across all currently input files, whileFNRit is the number of records in the current file. As in the following example:
$ awk '{print FILENAME, "FNR= ", FNR," NR= ", NR}' student-marks bookdetails
student-marks FNR= 1 NR= 1
student-marks FNR= 2 NR= 2
student-marks FNR= 3 NR= 3
student-marks FNR= 4 NR= 4
student-marks FNR= 5 NR= 5
bookdetails FNR= 1 NR= 6
bookdetails FNR= 2 NR= 7
bookdetails FNR= 3 NR= 8
bookdetails FNR= 4 NR= 9
bookdetails FNR= 5 NR= 10
Note: bookdetails and student-marks have the same content, used as an example. It can be seenNRandFNRthe difference.
Often usedNRandFNRin combination to process two files. For example, there are two files:
$ cat a.txt 李四|000002 张三|000001 王五|000003 赵六|000004 $ cat b.txt 000001|10 000001|20 000002|30 000002|15 000002|45 000003|40 000003|25 000004|60
If you want to make a correspondence, for example Zhang San|000001|10
$ awk -F '|' 'NR == FNR{a[$2]=$1;} NR>FNR {print a[$1],"|", $0}' a.txt b.txt
张三 | 000001|10
张三 | 000001|20
李四 | 000002|30
李四 | 000002|15
李四 | 000002|45
王五 | 000003|40
王五 | 000003|25
赵六 | 000004|60
English original: http://www.thegeekstuff.com/2010/01/8-powerful-awk-built-in-variables-fs-ofs-rs-ors-nr-nf-filename-fnr/
Translation: http://shomy.top/2016/05/05/trans-8-powerful-awk-built-in-variables/