BS4的快速使用

一、 核心前提与对象创建

  • 导入模块:安装包名为 beautifulsoup4,但代码中需导入 bs4 模块(from bs4 import BeautifulSoup)。
  • 解析不完整 HTML:Beautiful Soup 会自动补全缺失的标签(如 html、body 闭合标签)并修复结构。
  • 创建对象与格式化: soup = BeautifulSoup(html_doc, 'lxml') # 推荐显式指定 lxml 解析器 print(soup.prettify()) # 按标准缩进格式化输出完整结构

二、 结构化数据浏览(节点遍历与属性提取)

基于官方《爱丽丝梦游记》示例 HTML,常用快捷访问属性如下:

代码写法作用说明预期输出示例
soup.title获取第一个 <title> 标签对象<title>The Dormouse's story</title>
soup.title.name获取该标签的名称title
soup.title.string获取标签内的纯文本内容The Dormouse's story
soup.title.parent.name获取父级标签名称head
soup.p获取第一个 <p> 标签对象包含 class=”title” 的 p 标签
soup.p['class']获取标签的属性(返回列表)['title']
soup.a获取第一个 <a> 标签对象包含 Elsie 的 a 标签
soup.find_all('a')查找所有 <a> 标签(返回列表)包含 3 个 sister 的列表
soup.find(id="link3")按属性精确查找单个元素<a id="link3">Tillie</a>

三、 批量提取与全文获取

  • 提取所有链接: for link in soup.find_all('a'): print(link.get('href')) # 推荐使用 .get() 避免属性不存在时报错 # 输出:http://example.com/elsie 等 3 个链接
  • 获取全文纯文本: print(soup.get_text()) # 提取文档中所有文字,自动去除 HTML 标签

四、 完整可运行示例代码

from bs4 import BeautifulSoup

html_doc = """<html><head><title>The Dormouse's story</title></head>
<body>
<p class="title"><b>The Dormouse's story</b></p>
<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>
<p class="story">...</p>
</body></html>"""

soup = BeautifulSoup(html_doc, 'lxml')

# 1. 打印格式化结构
# print(soup.prettify())

# 2. 基础遍历演示
print("标题:", soup.title.string)
print("首个p的类名:", soup.p['class'])
print("按ID查找:", soup.find(id="link3").text)

# 3. 循环提取链接
for link in soup.find_all('a'):
    print("链接:", link.get('href'))

# 4. 提取全文
# print(soup.get_text())Code language: HTML, XML (xml)

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注