搜索

31-CSS搜索

按照CSS类名搜索tag的功能非常实用,但标识CSS类名的关键字 class 在Python中是保留字,使用 class 做参数会导致语法错误。从Beautiful Soup的4.1.1版本开始,可以通过 class_ 参数搜索有指定CSS类名的tag:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
all_tag = soup.find_all("a", class_="sister")
print(all_tag)Code language: HTML, XML (xml)

32-搜索页面中的文字text 参数

from bs4 import BeautifulSoup
import re

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie </a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find_all(text="Elsie "))
print(soup.find_all(text=re.compile("Dormouse")))Code language: HTML, XML (xml)

33-限制返回的个数limit参数

find_all() 方法返回全部的搜索结构,如果文档树很大那么搜索会很慢。如果我们不需要全部结果,可以使用 limit 参数限制返回结果的数量。效果与SQL中的limit关键字类似,当搜索到的结果数量达到 limit 的限制时,就停止搜索返回结果。

文档树中有3个tag符合搜索条件,但结果只返回了2个,因为我们限制了返回数量:

from bs4 import BeautifulSoup
import re

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie </a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find_all("a", limit=2))Code language: HTML, XML (xml)

34-只搜索tag的直接子节点recursive=False

调用tag的 find_all() 方法时,Beautiful Soup会检索当前tag的所有子孙节点,如果只想搜索tag的直接子节点,可以使用参数 recursive=False。

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie </a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.html.find_all("title"))
print(soup.html.find_all("title", recursive=False))Code language: HTML, XML (xml)

35-find查找一个

find_all() 方法将返回文档中符合条件的所有tag,尽管有时候我们只想得到一个结果。比如文档中只有一个 <body> 标签,那么使用 find_all() 方法来查找 <body> 标签就不太合适,使用 find_all 方法并设置 limit=1 参数不如直接使用 find() 方法。下面两行代码是等价的:

soup.find_all('title', limit=1)
# 等价于
soup.find('title')Code language: PHP (php)
from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie </a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find('p'))
print(soup.find_all('p', limit=1))Code language: HTML, XML (xml)

36-介绍一下其他的findxxx方法

find_parents() 和 find_parent()

find_parents(name, attrs, recursive, text, **kwargs)
find_parent(name, attrs, recursive, text, **kwargs)

find_all() 和 find() 只搜索当前节点的所有子节点、孙子节点等。find_parents() 和 find_parent() 用来搜索当前节点的父辈节点,搜索方法与普通tag的搜索方法相同。

a_string = soup.find(text="Lacie")
a_string
# 'Lacie'

a_string.find_parents("a")
# [<a class="sister" href="http://foxdevelop.com/lacie" id="link2">Lacie</a>]Code language: PHP (php)

find_next_siblings() 和 find_next_sibling()

find_next_siblings(name, attrs, recursive, text, **kwargs)
find_next_sibling(name, attrs, recursive, text, **kwargs)

这2个方法通过 .next_siblings 属性对当前tag的所有后面解析的兄弟tag节点进行迭代,find_next_siblings() 方法返回所有符合条件的后面的兄弟节点,find_next_sibling() 只返回符合条件的后面的第一个tag节点。

first_link = soup.a
first_link
# <a class="sister" href="http://foxdevelop.com/elsie" id="link1">Elsie</a>

first_link.find_next_siblings("a")
# [<a class="sister" href="http://foxdevelop.com/lacie" id="link2">Lacie</a>,
#  <a class="sister" href="http://foxdevelop.com/tillie" id="link3">Tillie</a>]

first_story_paragraph = soup.find("p", "story")
first_story_paragraph.find_next_sibling("p")
# <p class="story">...</p>Code language: HTML, XML (xml)

find_previous_siblings() 和 find_previous_sibling()

find_previous_siblings(name, attrs, recursive, text, **kwargs)
find_previous_sibling(name, attrs, recursive, text, **kwargs)

这2个方法通过 .previous_siblings 属性对当前tag的前面解析的兄弟tag节点进行迭代,find_previous_siblings() 方法返回所有符合条件的前面的兄弟节点,find_previous_sibling() 方法返回第一个符合条件的前面的兄弟节点。


37-css选择器

Beautiful Soup支持大部分的CSS选择器,在 Tag 或 BeautifulSoup 对象的 .select() 方法中传入字符串参数,即可使用CSS选择器的语法找到tag:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie </a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

# 找到某个tag标签下的直接子标签
print(soup.select("html > title"))
print(soup.select("p > a"))
print(soup.select("p > #link1"))

# 找到兄弟节点标签:
print(soup.select("#link1 ~ .sister"))

# 通过CSS的类名查找:
print(soup.select(".sister"))

# 通过tag的id查找:
print(soup.select("#link1"))
print(soup.select("a#link1"))

# 通过是否存在某个属性来查找:
print(soup.select('a[href]'))

# 通过属性的值来查找:
print(soup.select('a[href="http://foxdevelop.com/elsie"]'))Code language: HTML, XML (xml)

38-修改文档树

Beautiful Soup的强项是文档树的搜索,但同时也可以方便的修改文档树:

from bs4 import BeautifulSoup

soup = BeautifulSoup('<b class="boldest">Extremely bold</b>', "html5lib")

# 修改tag的名称和属性
tag = soup.b
tag.name = "blockquote"
tag['class'] = 'verybold'
tag['id'] = 1
print(tag)
# <blockquote class="verybold" id="1">Extremely bold</blockquote>

# 修改string
tag.string = "Foxdevelop"
print(tag)
# <blockquote class="verybold" id="1">Foxdevelop</blockquote>

# Tag.append() 方法向tag中添加内容,就好像Python的列表的 .append() 方法:
tag.append("www.foxdevelop.com")
print(tag)
# <blockquote class="verybold" id="1">Foxdevelopwww.foxdevelop.com</blockquote>

# 使用new_string来追加内容
new_string = soup.new_string(" 原创学习平台")
tag.append(new_string)
print(tag)
# <blockquote class="verybold" id="1">Foxdevelopwww.foxdevelop.com 原创学习平台</blockquote>Code language: PHP (php)

39-insert插入内容到指定的位置

Tag.insert() 方法与 Tag.append() 方法类似,区别是不会把新元素添加到父节点 .contents 属性的最后,而是把元素插入到指定的位置。与Python列表中的 .insert() 方法的用法相同:

from bs4 import BeautifulSoup

markup = '<a href="http://foxdevelop.com/">I linked to <i>example.com</i></a>'
soup = BeautifulSoup(markup, "html5lib")
tag = soup.a

tag.insert(1, "Foxdevelop ")
print(tag)
# <a href="http://foxdevelop.com/">I linked to Foxdevelop <i>example.com</i></a>

tag.insert(2, "Foxdevelop2 ")
print(tag)Code language: HTML, XML (xml)

replace_with()

PageElement.replace_with() 方法移除文档树中的某段内容,并用新tag或文本节点替代它:

markup = '<a href="http://foxdevelop.com/">I linked to <i>example.com</i></a>'
soup = BeautifulSoup(markup)
a_tag = soup.a

new_tag = soup.new_tag("b")
new_tag.string = "example.net"
a_tag.i.replace_with(new_tag)

a_tag
# <a href="http://foxdevelop.com/">I linked to <b>example.net</b></a>Code language: HTML, XML (xml)
Previous:

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注