获取内容

tag的 .contents 属性可以将tag的子节点以列表的方式输出

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>

<p class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

head_tag = soup.head

print(head_tag.contents)
print(head_tag.contents[0])
title_tag = head_tag.contents[0]
print(title_tag.contents)
print(title_tag.contents[0])

for child in title_tag.children:
    print(child)Code language: HTML, XML (xml)

.contents 和 .children 属性仅包含tag的直接子节点。例如,<head> 标签只有一个直接子节点 <title>,但是 <title> 标签也包含一个子节点:字符串 “The Dormouse’s story”,这种情况下字符串 “The Dormouse’s story” 也属于 <head> 标签的子孙节点。.descendants 属性可以对所有tag的子孙节点进行递归循环:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>

<p class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

head_tag = soup.head

for child in head_tag.descendants:
    print(child)

for child in head_tag.contents:
    print(child)

for child in head_tag.children:
    print(child)Code language: HTML, XML (xml)

如果tag只有一个 NavigableString 类型子节点,那么这个tag可以使用 .string 得到子节点。如果一个tag仅有一个子节点,那么这个tag也可以使用 .string 方法,输出结果与当前唯一子节点的 .string 结果相同。如果tag包含了多个子节点,tag就无法确定 .string 方法应该调用哪个子节点的内容,.string 的输出结果是 None:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<p class="title">www.foxdevelop.com</p>
<div class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</div>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

p_tag = soup.p
print(p_tag.string)

div_tag = soup.div
print(div_tag.string)
# noneCode language: HTML, XML (xml)

如果tag中包含多个字符串,可以使用 .strings 来循环获取。输出的字符串中可能包含了很多空格或空行,使用 .stripped_strings 可以去除多余空白内容:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<p class="title">www.foxdevelop.com</p>
<div class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</div>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

for string in soup.strings:
    print(repr(string))

for string in soup.stripped_strings:
    print(repr(string))Code language: HTML, XML (xml)

继续分析文档树,每个tag或字符串都有父节点:被包含在某个tag中。通过 .parent 属性来获取某个元素的父节点。在例子”爱丽丝”的文档中,<head> 标签是 <title> 标签的父节点:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<p class="title">www.foxdevelop.com</p>
<div class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</div>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

title_tag = soup.title
print(title_tag)
print(title_tag.parent)
print(title_tag.string.parent)
print(soup.html)
print(soup.html.parent)
print(soup.parent)Code language: HTML, XML (xml)

通过元素的 .parents 属性可以递归得到元素的所有父辈节点,下面的例子使用了 .parents 方法遍历了 <a> 标签到根节点的所有节点:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<p class="title">www.foxdevelop.com</p>
<div class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</div>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

link = soup.a
for parent in link.parents:
    if parent is None:
        print(parent)
    else:
        print(parent.name)Code language: HTML, XML (xml)

输出结果:

E:\桌面文件\untitled\venv\Scripts\python.exe E:/桌面文件/untitled/Demo.py
div
body
html
[document]Code language: JavaScript (javascript)

因为 <b> 标签和 <c> 标签是同一层:他们是同一个元素的子节点,所以 <b> 和 <c> 可以被称为兄弟节点。一段文档以标准格式输出时,兄弟节点有相同的缩进级别。在代码中也可以使用这种关系。虽然都是兄弟,但也区分哥哥和弟弟,所以下面我们来获取哥哥和弟弟两个节点。next是下一个,previous是上一个。在文档树中,使用 .next_sibling 和 .previous_sibling 属性来查询兄弟节点。<b> 标签有 .next_sibling 属性,但是没有 .previous_sibling 属性,因为 <b> 标签在同级节点中是第一个。同理,<c> 标签有 .previous_sibling 属性,却没有 .next_sibling 属性:

from bs4 import BeautifulSoup

sibling_soup = BeautifulSoup("<a><b>text1</b><c>text2</c></b></a>", 'html5lib')
print(sibling_soup.prettify())
print(sibling_soup.b.next_sibling)
print(sibling_soup.c.previous_sibling)

print(sibling_soup.b.previous_sibling)
# None
print(sibling_soup.c.next_sibling)
# None
print(sibling_soup.b.string.next_sibling)Code language: PHP (php)

通过 .next_siblings 和 .previous_siblings 属性可以对当前节点的兄弟节点迭代输出:

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<p class="title">www.foxdevelop.com</p>
<div class="title"><b>The Dormouse's story</b>sdfdsf<a></a>dsfdsfdsf</div>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')

for sibling in soup.a.next_siblings:
    print(repr(sibling))
for sibling in soup.find(id="link3").previous_siblings:
    print(repr(sibling))Code language: HTML, XML (xml)

28-搜索文档树

Beautiful Soup定义了很多搜索方法,这里着重介绍2个:find() 和 find_all()。其它方法的参数和用法类似,请读者举一反三。

再以”爱丽丝”文档作为例子:

过滤器

介绍 find_all() 方法前,先介绍一下过滤器的类型,这些过滤器贯穿整个搜索的API。过滤器可以被用在tag的name中、节点的属性中、字符串中或他们的混合中。

字符串

最简单的过滤器是字符串。在搜索方法中传入一个字符串参数,Beautiful Soup会查找与字符串完整匹配的内容,下面的例子用于查找文档中所有的 <b> 标签:

from bs4 import BeautifulSoup
import re

html_doc = """
<html><head><title>The Dormouse's story</title></head>

<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""
soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find_all('b'))Code language: HTML, XML (xml)

正则表达式

如果传入正则表达式作为参数,Beautiful Soup会通过正则表达式的 match() 来匹配内容。下面例子中找出所有以b开头的标签,这表示 <body> 和 <b> 标签都应该被找到:

for tag in soup.find_all(re.compile("^b")):
    print(tag.name)

# 下面代码找出所有名字中包含"t"的标签:
for tag in soup.find_all(re.compile("t")):
    print(tag.name)Code language: CSS (css)

29-find函数传入列表参数

如果传入列表参数,Beautiful Soup会将与列表中任一元素匹配的内容返回。下面代码找到文档中所有 <a> 标签和 <b> 标签:

from bs4 import BeautifulSoup
import re

html_doc = """
<html><head><title>The Dormouse's story</title></head>

<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""
soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find_all(["a", "b"]))

for tag in soup.find_all(True):
    print(tag.name)
# 等价于
for tag in soup.find_all():
    print(tag.name)Code language: HTML, XML (xml)

True 可以匹配任何值,下面代码查找到所有的tag,但是不会返回字符串节点。


30-自定义方法来实现过滤

如果没有合适过滤器,那么还可以定义一个方法,方法只接受一个元素参数,如果这个方法返回 True 表示当前元素匹配并且被找到,如果不是则返回 False。

下面方法校验了当前元素,如果包含 class 属性却不包含 id 属性,那么将返回 True:

keyword 参数

如果一个指定名字的参数不是搜索内置的参数名,搜索时会把该参数当作指定名字tag的属性来搜索,如果包含一个名字为 id 的参数,Beautiful Soup会搜索每个tag的”id”属性。

from bs4 import BeautifulSoup
import re

html_doc = """
<html><head><title>The Dormouse's story</title></head>

<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://foxdevelop.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://foxdevelop.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://foxdevelop.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""


def has_class_but_no_id(tag):
    return tag.has_attr('class') and not tag.has_attr('id')


# 将这个方法作为参数传入 find_all() 方法,将得到所有<p>标签:
soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.find_all(has_class_but_no_id))

print(soup.find_all(id="link2"))

# 使用多个指定名字的参数可以同时过滤tag的多个属性:
print(soup.find_all(href=re.compile("elsie"), id='link1'))Code language: HTML, XML (xml)

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注