一、 核心前提与对象创建
- 导入模块:安装包名为
beautifulsoup4,但代码中需导入bs4模块(from bs4 import BeautifulSoup)。 - 解析不完整 HTML:Beautiful Soup 会自动补全缺失的标签(如
html、body闭合标签)并修复结构。 - 创建对象与格式化:
soup = BeautifulSoup(html_doc, 'lxml') # 推荐显式指定 lxml 解析器 print(soup.prettify()) # 按标准缩进格式化输出完整结构
二、 结构化数据浏览(节点遍历与属性提取)
基于官方《爱丽丝梦游记》示例 HTML,常用快捷访问属性如下:
| 代码写法 | 作用说明 | 预期输出示例 |
|---|---|---|
soup.title | 获取第一个 <title> 标签对象 | <title>The Dormouse's story</title> |
soup.title.name | 获取该标签的名称 | title |
soup.title.string | 获取标签内的纯文本内容 | The Dormouse's story |
soup.title.parent.name | 获取父级标签名称 | head |
soup.p | 获取第一个 <p> 标签对象 | 包含 class=”title” 的 p 标签 |
soup.p['class'] | 获取标签的属性(返回列表) | ['title'] |
soup.a | 获取第一个 <a> 标签对象 | 包含 Elsie 的 a 标签 |
soup.find_all('a') | 查找所有 <a> 标签(返回列表) | 包含 3 个 sister 的列表 |
soup.find(id="link3") | 按属性精确查找单个元素 | <a id="link3">Tillie</a> |
三、 批量提取与全文获取
- 提取所有链接:
for link in soup.find_all('a'): print(link.get('href')) # 推荐使用 .get() 避免属性不存在时报错 # 输出:http://example.com/elsie 等 3 个链接 - 获取全文纯文本:
print(soup.get_text()) # 提取文档中所有文字,自动去除 HTML 标签
四、 完整可运行示例代码
from bs4 import BeautifulSoup
html_doc = """<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>
<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>
<p class="story">...</p>
</body></html>"""
soup = BeautifulSoup(html_doc, 'lxml')
# 1. 打印格式化结构
# print(soup.prettify())
# 2. 基础遍历演示
print("标题:", soup.title.string)
print("首个p的类名:", soup.p['class'])
print("按ID查找:", soup.find(id="link3").text)
# 3. 循环提取链接
for link in soup.find_all('a'):
print("链接:", link.get('href'))
# 4. 提取全文
# print(soup.get_text())Code language: HTML, XML (xml)
Previous: 什么是 Beautiful Soup
Next: BeautifulSoup对象