Python Using Regular Expressions to Extract URLs from Strings

Document 对象参考手册Python3 Examples

Given a string containing URL addresses, we need to use regular expressions to extract the URLs from the string.

Example

import re def Find(string): # findall() finds strings that match the regular expression url = re.findall('https?://(?:[-\w.]|(?:%[\da-fA-F]{2}))+', string) return url string = 'Example's web address is: https://www.example.com, Google's web address is: https://www.google.com' print("Urls: ", Find(string))

?:Explanation:

(?:x)

matchesxbut does not remember the match. This kind of parenthesis is called a non-capturing parenthesis, which allows you to define subexpressions to be used with regular expression operators. Look at this example/(?:foo){1,2}/. If the expression is/foo{1,2}/,{1,2}it will only apply to the last character 'o' of 'foo'. If non-capturing parentheses are used, then {1,2} will apply to the entire word 'foo'.

Executing the above code produces the following output:

Urls:  ['https://www.example.com', 'https://www.google.com']

Document 对象参考手册Python3 Examples

Other Extensions