Python爬虫之BeautifulSoup

作者: YungFan | 来源:发表于2017-06-14 09:48 被阅读101次

python爬虫之BeautifulSoup
BeautifulSoup requests 爬虫初体验
Python 爬虫
Python+PhantomJS+selenium+Beauti
男子大学生的無駄日常
Python爬虫之BeautifulSoup
bs4
Python爬虫入门（urllib+Beautifulsoup）
无标题文章
Python 爬虫实战（二）：使用 requests-html

上一篇博文中提到用正则表达式来匹配数据项，但是写起来容易出错，如果有过DOM开发经验或者使用过jQuery的朋友看到BeautifulSoup就像是见到了老朋友一样。

安装BeautifulSoup

Mac安装BeautifulSoup很简单，打开终端，执行以下语句，然后输入密码即可安装

sudo easy_install beautifulsoup4

改代码

#coding=utf-8
import urllib
from bs4 import BeautifulSoup

# 定义个函数 抓取网页内容
def getHtml(url):
    webPage = urllib.urlopen(url)
    html = webPage.read()
    return html

# 定义一个函数 抓取网页中的图片
def getNewsImgs(html):
    # 创建BeautifulSoup
    soup = BeautifulSoup(html, "html.parser")
    # 查找所有的img标签
    urlList = soup.find_all("img")
    length = len(urlList)
    # 遍历标签 下载图片
    for i in range(length):
        imgUrl =  urlList[i].attrs["src"]
        urllib.urlretrieve("http://www.abc.edu.cn/news/"+imgUrl,'news-%s.jpg' % i)

# 获取网页
html = getHtml("http://www.abc.edu.cn/news/show.aspx?id=21430&cid=5")
# 抓取图片
getNewsImgs(html)

效果：换了一个新闻，抓取了新闻中的三张图片_{O(∩_∩)O}~

爬虫抓图片.gif

网友评论

本文标题：Python爬虫之BeautifulSoup

本文链接：https://www.haomeiwen.com/subject/rhksqxtx.html

延伸阅读

深度阅读

您也可以注册成为美文阅读网的作者，发表您的原创作品、分享您的心情！

Python爬虫之BeautifulSoup

安装BeautifulSoup

改代码

效果：换了一个新闻，抓取了新闻中的三张图片_{O(∩_∩)O}~

相关文章

python爬虫之BeautifulSoup

BeautifulSoup requests 爬虫初体验

Python 爬虫

Python+PhantomJS+selenium+Beauti

男子大学生的無駄日常

Python爬虫之BeautifulSoup

bs4

Python爬虫入门（urllib+Beautifulsoup）

无标题文章

Python 爬虫实战（二）：使用 requests-html

网友评论

延伸阅读

深度阅读

栏目导航

热点阅读

首页投稿（暂停使用，暂停投稿）

程序员

Python爬虫之BeautifulSoup

安装BeautifulSoup

改代码

效果：换了一个新闻，抓取了新闻中的三张图片O(∩_∩)O~

相关文章

网友评论

延伸阅读

深度阅读

栏目导航

热点阅读

效果：换了一个新闻，抓取了新闻中的三张图片_{O(∩_∩)O}~