多个属性和标准html

10-多个属性值的问题

from bs4 import BeautifulSoup

soup = BeautifulSoup('<a></a><b class="boldest bai">中间内容</b>', 'html5lib')
print(soup.b)
print(soup.a)
print(soup.b["class"])
print(soup.b["class"][0])
print(soup.b["class"][1])
print(soup.b.get("class")[1])Code language: PHP (php)

输出结果:

<b class="boldest bai">中间内容</b>
<a></a>
['boldest', 'bai']
boldest
bai
baiCode language: HTML, XML (xml)

说明:class 属性在 HTML 模式下返回的是列表,可以通过索引或 get() 方法逐值访问。


11-格式化成标准HTML

有些时候,我们可能获得的 html 不是标准的 html 文档,比如缺少结束标记。这时候 BeautifulSoup 可以帮我们把 html 文档格式化为标准文档。

比如这样:

from bs4 import BeautifulSoup

soup = BeautifulSoup('<a></a><b class="boldest bai">中间内容</b>', 'html5lib')
print(soup.html)Code language: HTML, XML (xml)

本身是没有 html 标签的,但是我们却可以输出 html 对象。

最后输出:

<html><head></head><body><a></a><b class="boldest bai">中间内容</b></body></html>Code language: HTML, XML (xml)

就算我们的 html 是不完整的,比如缺少了结束标记,也可以正常解析:

from bs4 import BeautifulSoup

soup = BeautifulSoup('<a><b class="boldest bai">中间内容<b>', 'html5lib')
print(soup)Code language: HTML, XML (xml)

12-一个例子

我们看一个例子,例子中有一段不完整的 html,bs 会帮我们把不完整的文档补充完整。然后我们来获取 title 以及 title 的内容,还可以获取多个,获取标签中的属性。

from bs4 import BeautifulSoup

html_doc = """
<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>

<p class="story">...</p>
"""

soup = BeautifulSoup(html_doc, 'html5lib')
print(soup.title)
# <title>The Dormouse's story</title>

print(soup.title.name)
# title

print(soup.title.string)
# The Dormouse's story

print(soup.title.parent.name)
# head

print(soup.p)
# <p class="title"><b>The Dormouse's story</b></p>

print(soup.p['class'])
# ['title']

print(soup.a)
# <a class="sister" href="http://example.com/elsie" id="link1">Elsie</a>

print(soup.find_all('a'))
# [<a class="sister" href="http://example.com/elsie" id="link1">Elsie</a>,
#  <a class="sister" href="http://example.com/lacie" id="link2">Lacie</a>,
#  <a class="sister" href="http://example.com/tillie" id="link3">Tillie</a>]

print(soup.find(id="link3"))
# <a class="sister" href="http://example.com/tillie" id="link3">Tillie</a>Code language: HTML, XML (xml)

从文档中找到所有 <a> 标签的链接:

for link in soup.find_all('a'):
    print(link.get('href'))
    # http://example.com/elsie
    # http://example.com/lacie
    # http://example.com/tillieCode language: PHP (php)

从文档中获取所有文字内容:

print(soup.get_text())
# The Dormouse's story
#
# The Dormouse's story
#
# Once upon a time there were three little sisters; and their names were
# Elsie,
# Lacie and
# Tillie;
# and they lived at the bottom of a well.
#
# ...Code language: PHP (php)

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注