爬虫_043_beautifulsoup的css选择器

作者: 为宇绸缪 | 来源:发表于2023-03-05 19:00 被阅读0次

python-爬虫
爬虫课程大纲
GO学习笔记(25) - 爬虫(2) - 选择器等工具
CSS选择器
写个爬虫
2-BeautifulSoup4
pyspider 爬虫教程
CSS选择器
css选择器
CSS 选择器

我们在写 CSS 时，标签名不加任何修饰，类名前加点，id名前加 #，在这里我们也可以利用类似的方法来筛选元素，用到的方法是 soup.select()，返回类型是 list

1、通过标签名查找

print(soup.select("title"))  #[<title>The Dormouse's story</title>]
print(soup.select("b"))      #[<b>The Dormouse's story</b>]

2、通过类名查找

print(soup.select(".sister")) 

'''
[<a class="sister" href="http://example.com/elsie" id="link1">Elsie</a>, 
<a class="sister" href="http://example.com/lacie" id="link2">Lacie</a>, 
<a class="sister" href="http://example.com/tillie" id="link3">Tillie</a>]

'''

3、id名查找

print(soup.select("#link1"))
# [<a class="sister" href="http://example.com/elsie" id="link1">Elsie</a>]

4、组合查找

组合查找即和写 class 文件时，标签名与类名、id名进行的组合原理是一样的，例如查找 p 标签中，id 等于 link1的内容，二者需要用空格分开

print(soup.select("p #link2"))

#[<a class="sister" href="http://example.com/lacie" id="link2">Lacie</a>]

直接子标签查找

print(soup.select("p > #link2"))
# [<a class="sister" href="http://example.com/lacie" id="link2">Lacie</a>]

查找既有class也有id选择器的标签

a_string = soup.select(".story#test")

查找有多个class选择器的标签

a_string = soup.select(".story.test")

查找有多个class选择器和一个id选择器的标签

a_string = soup.select(".story.test#book")

5、属性查找

查找时还可以加入属性元素，属性需要用中括号括起来，注意属性和标签属于同一节点，所以中间不能加空格，否则会无法匹配到。

print(soup.select("a[href='http://example.com/tillie']"))
#[<a class="sister" href="http://example.com/tillie" id="link3">Tillie</a>]

select 方法返回的结果都是列表形式，可以遍历形式输出，然后用 get_text() 方法来获取它的内容：

for title in soup.select('a'):
    print (title.get_text())

'''
Elsie
Lacie
Tillie
'''

网友评论

js css html

本文标题：爬虫_043_beautifulsoup的css选择器

本文链接：https://www.haomeiwen.com/subject/kmhpldtx.html

延伸阅读

深度阅读

您也可以注册成为美文阅读网的作者，发表您的原创作品、分享您的心情！

爬虫_043_beautifulsoup的css选择器

1、通过标签名查找

2、通过类名查找

3、id名查找

4、组合查找

5、属性查找

相关文章

python-爬虫

爬虫课程大纲

GO学习笔记(25) - 爬虫(2) - 选择器等工具

CSS选择器

写个爬虫

2-BeautifulSoup4

pyspider 爬虫教程

CSS选择器

css选择器

CSS 选择器

网友评论

延伸阅读

深度阅读

栏目导航

热点阅读

js css html