Ruby String

In Ruby, String objects are used to store or manipulate sequences of one or more bytes.

Ruby strings are divided into single-quoted strings (') and double-quoted strings ("). The difference is that double-quoted strings support more escape characters.

Single-quoted strings

The simplest strings are single-quoted strings, i.e., strings enclosed within single quotes:

'This is a string in a Ruby program'

If you need to use a single-quote character inside a single-quoted string, you need to use a backslash (\) within the single-quoted string, so that the Ruby interpreter will not treat that single-quote character as the terminator of the string:

'Won\'t you read O\'Reilly\'s book?'

A backslash can also escape another backslash, so that the second backslash itself is not interpreted as an escape character.

The following are features related to strings in Ruby.

Double-quoted strings

In double-quoted strings, we can use#{}the number sign and curly braces to evaluate expressions:

Embedding variables in strings:

Examples

#!/usr/bin/ruby # -*- coding: UTF-8 -*- name1 = "Joe" name2 = "Mary" puts "Hello #{name1}, where is #{name2}?"

Try it »

The above example outputs the following result:

你好 Joe,  Mary 在哪?

Performing mathematical operations in strings:

Examples

#!/usr/bin/ruby # -*- coding: UTF-8 -*- x, y, z = 12, 36, 72 puts "The value of x is #{ x }" puts "The value of x + y is #{ x + y }" puts "The average of x + y + z is #{ (x + y + z)/3 }"

Try it »

The above example outputs the following result:

x 的值为 12
x + y 的值为 48
x + y + z 的平均值为 40

Ruby also supports string literals introduced by %q and %Q. %q uses single-quote quoting rules, while %Q uses double-quote quoting rules, followed by a start delimiter such as ( ! [ { etc., and a corresponding end delimiter such as } ] ).

The character following q or Q is the delimiter. The delimiter can be any non-alphanumeric single-byte character, such as [, {, (, <, !, etc. The string is read until a matching closing delimiter is found.

Examples

#!/usr/bin/ruby # -*- coding: UTF-8 -*- desc1 = %Q{Ruby strings can use '' and "".} desc2 = %q|Ruby strings can use '' and "".| puts desc1 puts desc2

Try it »

The above example outputs the following result:

Ruby 的字符串可以使用 '' 和 ""。
Ruby 的字符串可以使用 '' 和 ""。

Escape characters

The table below lists escape characters or non-printable characters that can be escaped using the backslash notation.

Note:Inside a double-quoted string, escape characters are parsed. Inside a single-quoted string, escape characters are not parsed and are output as-is.

Backslash notationHexadecimal characterDescription
\a0x07Alert (bell)
\b0x08Backspace
\cx Control-x
\C-x Control-x
\e0x1bEscape
\f0x0cForm feed
\M-\C-x Meta-Control-x
\n0x0aNewline
\nnn Octal representation, where n ranges from 0 to 7
\r0x0dCarriage return
\s0x20Space
\t0x09Tab
\v0x0bVertical tab
\x Character x
\xnn Hexadecimal representation, where n ranges from 0-9, a-f, or A-F

Character encoding

Ruby's default character set is ASCII, where characters can be represented by a single byte. If you use UTF-8 or another modern character set, characters may be represented by one to four bytes.

You can use $KCODE at the beginning of the program to change the character set, as shown below:

$KCODE = 'u'

The following are the possible values of $KCODE.

EncodingDescription
aASCII (same as none). This is the default.
eEUC。
nNone (same as ASCII).
uUTF-8。

Built-in string methods

We need an instance of the String object to call String methods. The following is how to create an instance of a String object:

new [String.new(str="")]

This will return a new string object containingstra copy. Now, usingstrthe object, we can call any available instance method. For example:

Examples

#!/usr/bin/ruby myStr = String.new("THIS IS TEST") foo = myStr.downcase puts "#{foo}"

This will produce the following result:

this is test

The following are the public string methods (assuming str is a String object):

No.Method & Description
1str % arg
Formats a string using a format specification. If arg contains more than one substitution, then arg must be an array. For more information on format specifications, see sprintf under the Kernel module.
2str * integer
Returns a new string containing integer copies of str. In other words, str is repeated integer times.
3str + other_str
Concatenates other_str to str.
4str << obj
Concatenates an object to the string. If the object is a Fixnum with a value between 0 and 255, it is converted to a character. Compare this with concat.
5str <=> other_str
Compares str with other_str, returning -1 (less than), 0 (equal to), or 1 (greater than). The comparison is case-sensitive.
6str == obj
Checks the equality of str and obj. If obj is not a string, returns false; if str <=> obj returns 0, returns true.
7str =~ obj
Matches str against the regular expression pattern obj. Returns the position where the match starts, otherwise returns false.
8str[position] # Note: returns the ASCII code rather than a character
str[start, length]
str[start..end]
str[start...end]

Extracts a substring using an index.
9str.capitalize
Converts the first letter of a string to uppercase and the rest to lowercase for display.
10str.capitalize!
Same as capitalize, but capitalize! returns nil if no changes are made.
11str.casecmp
Case-insensitive string comparison.
12str.center
Centers the string.
13str.chomp
Removes the record separator ($/) from the end of the string, usually \n. If there is no record separator, it does nothing.
14str.chomp!
Same as chomp, but str is modified and returned.
15str.chop
Removes the last character from str.
16str.chop!
Same as chop, but str is modified and returned.
17str.concat(other_str)
Concatenates other_str to str.
18str.count(str, ...)
Counts one or more character sets. If there are multiple character sets, it counts the intersection of these sets.
19str.crypt(other_str)
Applies a one-way cryptographic hash to str. The argument is a two-character-long string, with each character in the range a-z, A-Z, 0-9, ., or /.
20str.delete(other_str, ...)
Returns a copy of str with all characters in the intersection of the arguments deleted.
21str.delete!(other_str, ...)
Same as delete, but str is modified and returned.
22str.downcase
Returns a copy of str with all uppercase letters replaced by lowercase letters.
23str.downcase!
Same as downcase, but str is modified and returned.
24str.dump
Returns a version of str with all non-printable characters replaced by \nnn notation and all special characters escaped.
25str.each(separator=$/) { |substr| block }
Divides str using the argument as the record separator (default is $/), passing each substring to the supplied block.
26str.each_byte { |fixnum| block }
Passes each byte of str to the block, returning each byte in decimal notation.
27str.each_line(separator=$/) { |substr| block }
Divides str using the argument as the record separator (default is $/), passing each substring to the supplied block.
28str.empty?
Returns true if str is empty (i.e., has a length of 0).
29str.eql?(other)
Two strings are equal if they have the same length and content.
30str.gsub(pattern, replacement) [or]
str.gsub(pattern) { |match| block }

Returns a copy of str with all occurrences of pattern replaced by replacement or the value of block. pattern is usually a regexp Regexp; if it is a String, no regular expression metacharacters are interpreted (i.e., /\d/ will match a digit, but '\d' will match a backslash followed by a 'd').
31str[fixnum] [or] str[fixnum,fixnum] [or] str[range] [or] str[regexp] [or] str[regexp, fixnum] [or] str[other_str]
Use the following parameter references for str: if the parameter is a Fixnum, return the character encoding of the fixnum; if the parameter is two Fixnums, return a substring starting at the offset (the first fixnum) up to the length (the second fixnum); if the parameter is a range, return a substring within that range; if the parameter is a regexp, return the part of the string that matches; if the parameter is a regexp with a fixnum, return the match data at the fixnum position; if the parameter is other_str, return the substring that matches other_str. A negative Fixnum starts from the end of the string at -1.
32str[fixnum] = fixnum [or] str[fixnum] = new_str [or] str[fixnum, fixnum] = new_str [or] str[range] = aString [or] str[regexp] =new_str [or] str[regexp, fixnum] =new_str [or] str[other_str] = new_str ]
Replace the entire string or part of the string. Synonym for slice!.
33str.gsub!(pattern, replacement) [or] str.gsub!(pattern) { |match| block }
Perform the replacements of String#gsub, returning str, or return nil if no replacement was performed.
34str.hash
Return a hash based on the length and content of the string.
35str.hex
Treat the leading characters of str as a string of hexadecimal digits (an optional sign and an optional 0x), and return the corresponding number. Return zero on error.
36str.include? other_str [or] str.include? fixnum
Return true if str contains the given string or character.
37str.index(substring [, offset]) [or]
str.index(fixnum [, offset]) [or]
str.index(regexp [, offset])

Return the index of the first occurrence of the given substring, character (fixnum), or pattern (regexp) in str. Return nil if not found. If a second parameter is provided, it specifies the position in the string to begin the search.
38str.insert(index, other_str)
Insert other_str before the character at the given index, modifying str. Negative index values count from the end of the string and insert after the given character. The intent is to start inserting a string at the given index.
39str.inspect
Return a printable version of str, with special characters escaped.
40str.intern [or] str.to_sym
Return the symbol corresponding to str, creating the symbol if it does not previously exist.
41str.length
Return the length of str. Compare it with size.
42str.ljust(integer, padstr=' ')
If integer is greater than the length of str, return a new string of length integer with str left-aligned and padded with padstr. Otherwise, return str.
43str.lstrip
Return a copy of str with leading whitespace removed.
44str.lstrip!
Remove leading whitespace from str, or return nil if there is no change.
45str.match(pattern)
If pattern is not a regular expression, convert pattern to a Regexp, then call its match method on str.
46str.oct
Treat the leading characters of str as a string of decimal digits (an optional sign), and return the corresponding number. Return 0 if the conversion fails.
47str.replace(other_str)
Replace the contents of str with the corresponding values from other_str.
48str.reverse
Return a new string that is the reverse of str.
49str.reverse!
Reverse str, changing str in place and returning it.
50str.rindex(substring [, fixnum]) [or]
str.rindex(fixnum [, fixnum]) [or]
str.rindex(regexp [, fixnum])

Return the index of the last occurrence of the given substring, character (fixnum), or pattern (regexp) in str. Return nil if not found. If a second parameter is provided, it specifies the position in the string to end the search. Characters beyond that point are not considered.
51str.rjust(integer, padstr=' ')
If integer is greater than the length of str, return a new string of length integer with str right-aligned and padded with padstr. Otherwise, return str.
52str.rstrip
Return a copy of str with trailing whitespace removed.
53str.rstrip!
Remove trailing whitespace from str, or return nil if there is no change.
54str.scan(pattern) [or]
str.scan(pattern) { |match, ...| block }

Both forms match pattern (which can be a Regexp or a String) against str. For each match, a result is generated and added to the result array or passed to the block. If pattern contains no groups, each individual result consists of the matched string, $&. If pattern contains groups, each individual result is an array containing each group entry.
55str.slice(fixnum) [or] str.slice(fixnum, fixnum) [or]
str.slice(range) [or] str.slice(regexp) [or]
str.slice(regexp, fixnum) [or] str.slice(other_str)
See str[fixnum], etc.
str.slice!(fixnum) [or] str.slice!(fixnum, fixnum) [or] str.slice!(range) [or] str.slice!(regexp) [or] str.slice!(other_str)

Delete the specified portion from str and return the deleted portion. If the value is out of range, the form with a Fixnum parameter generates an IndexError. The range form generates a RangeError, and the Regexp and String forms ignore the action.
56str.split(pattern=$;, [limit])

Split str into substrings based on a delimiter, and return an array of these substrings.

Ifpatternis a String, then when splitting str, it is used as the delimiter. If pattern is a single space, then str is split based on whitespace, ignoring leading whitespace and consecutive whitespace characters.

Ifpatternis a Regexp, then str is split at places where pattern matches. When pattern matches a zero-length string, str is split into individual characters.

If thepatternparameter is omitted, the value of $; is used. If $; is nil (the default), str is split based on whitespace, as if ` ` were specified as the delimiter.

If thelimitparameter is omitted, trailing null fields are suppressed. If limit is a positive number, at most that number of fields is returned (if limit is 1, the entire string is returned as the only entry in the array). If limit is a negative number, the number of fields returned is unlimited, and trailing null fields are not suppressed.

57str.squeeze([other_str]*)
Build a set of characters from the other_str parameters using the procedure described for String#count. Return a new string where identical characters appearing in the set are replaced with a single character. If no arguments are given, all identical characters are replaced with a single character.
58str.squeeze!([other_str]*)
Same as squeeze, but str is changed in place and returned, or nil if no change.
59str.strip
Return a copy of str with leading and trailing whitespace removed.
60str.strip!
Remove leading and trailing whitespace from str, or return nil if there is no change.
61str.sub(pattern, replacement) [or]
str.sub(pattern) { |match| block }

Return a copy of str with the first occurrence of pattern replaced by replacement or the value of the block. pattern is usually a Regexp; if it is a String, no regular expression metacharacters are interpreted.
62str.sub!(pattern, replacement) [or]
str.sub!(pattern) { |match| block }

Perform the substitutions of String#sub, and return str, or return nil if no substitution was performed.
63str.succ [or] str.next
Return the successor of str.
64str.succ! [or] str.next!
Equivalent to String#succ, but str is changed in place and returned.
65str.sum(n=16)
Return the n-bit checksum of the characters in str, where n is the optional Fixnum parameter, defaulting to 16. The result is simply the sum of the binary values of each character in str, modulo 2n - 1. This is not a particularly good checksum.
66str.swapcase
Return a copy of str with all uppercase letters converted to lowercase and all lowercase letters converted to uppercase.
67str.swapcase!
Equivalent to String#swapcase, but str is changed in place and returned, or nil if no change.
68str.to_f
Return the result of interpreting the leading characters of str as a floating point number. Extra characters beyond the end of the significant digits are ignored. If there is no valid number at the beginning of str, return 0.0. This method does not generate exceptions.
69str.to_i(base=10)
Return the result of interpreting the leading characters of str as an integer base (base 2, 8, 10, or 16). Extra characters beyond the end of the valid digits are ignored. If there is no valid number at the beginning of str, return 0. This method does not generate exceptions.
70str.to_s [or] str.to_str
Return the receiver.
71str.tr(from_str, to_str)
Return a copy of str with the characters in from_str replaced by the corresponding characters in to_str. If to_str is shorter than from_str, it is padded with its last character. Both strings can use the c1.c2 notation to represent ranges of characters. If from_str begins with ^, it means all characters except those listed.
72str.tr!(from_str, to_str)
Equivalent to String#tr, but str is changed in place and returned, or nil if no change.
73str.tr_s(from_str, to_str)
Process str according to the rules described for String#tr, then remove duplicate characters that would affect the translation.
74str.tr_s!(from_str, to_str)
Equivalent to String#tr_s, but str is changed in place and returned, or nil if no change.
75str.unpack(format)
Decode str (which may contain binary data) according to the format string, returning an array of each extracted value. The format characters consist of a series of single-character directives. Each directive can be followed by a number, indicating the number of times to repeat that directive. An asterisk (*) uses all remaining elements. The directives sSiIlL may each be followed by an underscore (_) to use the native size of the underlying platform for the specified type; otherwise, a platform-independent consistent size is used. Whitespace in the format string is ignored.
76str.upcase
Return a copy of str with all lowercase letters replaced by uppercase letters. The operation is environment-insensitive; only the characters a to z are affected.
77str.upcase!
Change the content of str to uppercase, or return nil if there is no change.
78str.upto(other_str) { |s| block }
Iterate over consecutive values, starting with str and ending with other_str (inclusive), passing each value in turn to the block. The String#succ method is used to generate each value.

String unpack directives

The following table lists the unpack directives for the method String#unpack.

InstructionReturnDescription
AStringRemove trailing nulls and spaces.
aStringString.
BStringExtract bits from each character (most significant bit first).
bStringExtract bits from each character (least significant bit first).
CFixnumExtract a character as an unsigned integer.
cFixnumExtract a character as an integer.
D, dFloatTreat sizeof(double) characters as a native double.
EFloatTreat sizeof(double) characters as a littleendian double.
eFloatTreat sizeof(float) characters as a littleendian float.
F, fFloatTreat sizeof(float) characters as a native float.
GFloatTreat sizeof(double) characters as a network byte order double.
gFloatTreat sizeof(float) characters as a network byte order float.
HStringExtract hexadecimal from each character (most significant first).
hStringExtract hexadecimal from each character (least significant first).
IIntegerTreat sizeof(int) length (modified by _) consecutive characters as a native integer.
iIntegerTreat sizeof(int) length (modified by _) consecutive characters as a signed native integer.
LIntegerTreat four (modified by _) consecutive characters as an unsigned native long integer.
lIntegerTreat four (modified by _) consecutive characters as a signed native long integer.
MStringQuoted-printable.
mStringBase64 encoding.
NIntegerTreat four characters as an unsigned long in network byte order.
nFixnumTreat two characters as an unsigned short in network byte order.
PStringTreat sizeof(char *) characters as a pointer, and return \emph{len} characters from the referenced location.
pStringTreat sizeof(char *) characters as a pointer to a null-terminated character string.
QIntegerTreat eight characters as an unsigned quad word (64-bit).
qIntegerTreat eight characters as a signed quad word (64-bit).
SFixnumTreat two consecutive characters (different if _ is used) as an unsigned short in native byte order.
sFixnumTreat two consecutive characters (different if _ is used) as a signed short in native byte order.
UIntegerUTF-8 character as an unsigned integer.
uStringUU encoding.
VFixnumTreat four characters as an unsigned long in little-endian byte order.
vFixnumTreat two characters as an unsigned short in little-endian byte order.
wIntegerBER-compressed integer.
X Skip one character backward.
x Skip one character forward.
ZStringUsed with *, remove trailing nulls up to the first null.
@ Skip the offset given by the length parameter.

Examples

Try the following examples to unpack various data.

"abc \0\0abc \0\0".unpack('A6Z6') #=> ["abc", "abc "] "abc \0\0".unpack('a3a3') #=> ["abc", " \000\000"] "abc \0abc \0".unpack('Z*Z*') #=> ["abc ", "abc "] "aa".unpack('b8B8') #=> ["10000110", "01100001"] "aaa".unpack('h2H2c') #=> ["16", "61", 97] "\xfe\xff\xfe\xff".unpack('sS') #=> [-2, 65534] "now=20is".unpack('M*') #=> ["now is"] "whole".unpack('xax2aX2aX1aX2a') #=> ["h", "e", "l", "l", "o"]
Other extensions