13-Beautiful Soup的各种对象
Beautiful Soup将复杂HTML文档转换成一个复杂的树形结构,每个节点都是Python对象,所有对象可以归纳为4种:Tag、NavigableString、BeautifulSoup、Comment。
1 Tag对象
Tag 对象与XML或HTML原生文档中的tag相同:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag)
print(type(tag))
# <class 'bs4.element.Tag'>
tag = soup.a
print(tag)
# <a class="1a"></a>
print(type(tag))
# <class 'bs4.element.Tag'>Code language: HTML, XML (xml)
BS会自动查找获取第一个,所以上面有两个a标签,取的是第一个a标签。
2 Name
获取标签名称:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag.name)
# bCode language: HTML, XML (xml)
如果改变了tag的name,那将影响所有通过当前Beautiful Soup对象生成的HTML文档:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag.name)
# b
tag.name = "blockquote"
print(tag)
# <blockquote class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</blockquote>Code language: HTML, XML (xml)
3 Attributes
一个tag可能有很多个属性。tag <b class="boldest"> 有一个 “class” 的属性,值为 “boldest”。tag的属性的操作方法与字典相同:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag["class"])
# ['boldest']Code language: HTML, XML (xml)
tag的属性可以被添加、删除或修改。再说一次,tag的属性操作方法与字典一样:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag["class"])
# ['boldest']
tag["class"] = "abc"
print(tag)
# <b class="abc"><a class="1a"></a><a class="2a"></a>Extremely bold</b>
tag["id"] = 1
print(tag)
# <b class="abc" id="1"><a class="1a"></a><a class="2a"></a>Extremely bold</b>
del tag["class"]
print(tag)
# <b id="1"><a class="1a"></a><a class="2a"></a>Extremely bold</b>Code language: HTML, XML (xml)
from bs4 import BeautifulSoup
soup = BeautifulSoup('<b class="boldest dsf"><a class="1a"></a><a class="2a"></a>Extremely bold</b>', 'html5lib')
tag = soup.b
print(tag)
print(type(tag))
# <class 'bs4.element.Tag'>
print(tag.name)
tag.name = "div"
print(tag)
print(tag["class"][0])
tag["class"][0] = "fast"
print(tag)
del tag["class"][1]
print(tag)Code language: PHP (php)
14-一个属性多个值的情况
HTML 4定义了一系列可以包含多个值的属性。在HTML5中移除了一些,却增加更多。最常见的多值的属性是 class(一个tag可以有多个CSS的class)。还有一些属性 rel、rev、accept-charset、headers、accesskey。在Beautiful Soup中多值属性的返回类型是list:
from bs4 import BeautifulSoup
css_soup = BeautifulSoup('<p class="body strikeout"></p>')
print(css_soup.p['class'])
# ["body", "strikeout"]
css_soup = BeautifulSoup('<p class="body"></p>')
print(css_soup.p['class'])
# ["body"]Code language: PHP (php)
如果某个属性看起来好像有多个值,但在任何版本的HTML定义中都没有被定义为多值属性,那么Beautiful Soup会将这个属性作为字符串返回:
from bs4 import BeautifulSoup
id_soup = BeautifulSoup('<p id="my id"></p>')
print(id_soup.p['id'])
# my idCode language: PHP (php)
可以给多值属性赋值,赋值的值是数组:
from bs4 import BeautifulSoup
rel_soup = BeautifulSoup('<p>Back to the <a rel="index">homepage</a></p>')
print(rel_soup.a['rel'])
# ['index']
rel_soup.a['rel'] = ['index', 'contents']
print(rel_soup.p)
# <p>Back to the <a rel="index contents">homepage</a></p>Code language: PHP (php)
如果转换的文档是XML格式,那么tag中不包含多值属性:
from bs4 import BeautifulSoup
xml_soup = BeautifulSoup('<p class="body strikeout"></p>', 'xml')
print(xml_soup.p['class'])Code language: JavaScript (javascript)
安装xml:
C:\Users\Administrator\PycharmProjects\untitled\venv\Scripts\python.exe C:/Users/Administrator/PycharmProjects/untitled/Demo.py
Traceback (most recent call last):
File "C:/Users/Administrator/PycharmProjects/untitled/Demo.py", line 3, in <module>
xml_soup = BeautifulSoup('<p class="body strikeout"></p>', 'xml')
File "C:\Users\Administrator\PycharmProjects\untitled\venv\lib\site-packages\bs4\__init__.py", line 196, in __init__
% ",".join(features))
bs4.FeatureNotFound: Couldn't find a tree builder with the features you requested: xml. Do you need to install a parser library?
Process finished with exit code 1Code language: PHP (php)
15-可以遍历的字符串
字符串常被包含在tag内。Beautiful Soup用 NavigableString 类来包装tag中的字符串:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<p class="body strikeout">www.foxdevelop.com Foxdevelop</p>', 'html5lib')
print(soup.p.string)
print(type(soup.p.string))
# <class 'bs4.element.NavigableString'>Code language: PHP (php)
一个 NavigableString 字符串与Python中的Unicode字符串相同,并且还支持包含在遍历文档树和搜索文档树中的一些特性。通过 unicode() 方法可以直接将 NavigableString 对象转换成Unicode字符串:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<p class="body strikeout">www.foxdevelop.com Foxdevelop</p>', 'html5lib')
print(soup.p.string)
print(type(soup.p.string))
# <class 'bs4.element.NavigableString'>
soup.p.string.replace_with("No longer bold")
print(soup.p)Code language: PHP (php)
16-BeautifulSoup对象
BeautifulSoup 对象表示的是一个文档的全部内容。大部分时候,可以把它当作 Tag 对象,它支持遍历文档树和搜索文档树中描述的大部分的方法。
因为 BeautifulSoup 对象并不是真正的HTML或XML的tag,所以它没有name和attribute属性。但有时查看它的 .name 属性是很方便的,所以 BeautifulSoup 对象包含了一个值为 “[document]” 的特殊属性 .name:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<p class="body strikeout">www.foxdevelop.com Foxdevelop</p>', 'html5lib')
print(soup)
print(type(soup))
# <class 'bs4.BeautifulSoup'>
print(soup.name)
# [document]Code language: PHP (php)
17-注释及特殊字符串
Tag、NavigableString、BeautifulSoup 几乎覆盖了html和xml中的所有内容,但是还有一些特殊对象。容易让人担心的内容是文档的注释部分:
from bs4 import BeautifulSoup
markup = "<b><!--Hey, buddy. Want to buy a used parser?--></b>"
soup = BeautifulSoup(markup, 'html5lib')
comment = soup.b.string
print(type(comment))
# <class 'bs4.element.Comment'>Code language: HTML, XML (xml)
就算注释不完整,漏掉后面的 > 也可以正常解析。Comment 对象是一个特殊类型的 NavigableString 对象:
from bs4 import BeautifulSoup
markup = "<b><!--Hey, buddy. Want to buy a used parser?--></b>"
soup = BeautifulSoup(markup, 'html5lib')
comment = soup.b.string
print(type(comment))
# <class 'bs4.element.Comment'>
print(comment)
print(soup.b.prettify())Code language: HTML, XML (xml)