Python爬取酷狗音乐TOP500 !

作者: 14e61d025165 | 来源:发表于2019-04-26 15:44 被阅读0次

python第四天（一）BeautifulSoup爬虫
Python爬取酷狗音乐TOP500 !
Python爬取酷狗音乐
python爬虫教程：爬取酷狗音乐
python爬虫教程：爬取酷狗音乐！
Python爬取酷狗音乐py文件
爬取酷狗TOP500的数据
用Python爬取酷我音乐
Java爬取并下载酷狗TOP500歌曲
流量时代，何去何从

发这个主要是因为我本地没有歌，有的歌还是VIP下载不了，平时听歌还得用流量。所以就想着看能直接把所有的歌曲直接拿下来。就去看了酷狗的主页面。想直接拿到TOP500.因为没找到怎么去下载，然后就在网上找了一下，找到了一个根据hash拼接url，下载歌曲。，只要找到hash值就啥都解决了。

看了之后发现还不是特别难弄，所以就直接趴下来了。

<tt-image data-tteditor-tag="tteditorTag" contenteditable="false" class="syl1556264619765 ql-align-center" data-render-status="finished" data-syl-blot="image" style="box-sizing: border-box; cursor: text; text-align: left; color: rgb(34, 34, 34); font-family: "PingFang SC", "Hiragino Sans GB", "Microsoft YaHei", "WenQuanYi Micro Hei", "Helvetica Neue", Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; white-space: pre-wrap; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; background-color: rgb(255, 255, 255); text-decoration-style: initial; text-decoration-color: initial; display: block;">

image

上面是网址，

<tt-image data-tteditor-tag="tteditorTag" contenteditable="false" class="syl1556264619770 ql-align-center" data-render-status="finished" data-syl-blot="image" style="box-sizing: border-box; cursor: text; text-align: left; color: rgb(34, 34, 34); font-family: "PingFang SC", "Hiragino Sans GB", "Microsoft YaHei", "WenQuanYi Micro Hei", "Helvetica Neue", Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; white-space: pre-wrap; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; background-color: rgb(255, 255, 255); text-decoration-style: initial; text-decoration-color: initial; display: block;">

image

改变数字就可以实现翻页，所以这个不能翻页的问题解决了。然后就是老套路按F12查看找network.

<tt-image data-tteditor-tag="tteditorTag" contenteditable="false" class="syl1556264619772 ql-align-center" data-render-status="finished" data-syl-blot="image" style="box-sizing: border-box; cursor: text; text-align: left; color: rgb(34, 34, 34); font-family: "PingFang SC", "Hiragino Sans GB", "Microsoft YaHei", "WenQuanYi Micro Hei", "Helvetica Neue", Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; white-space: pre-wrap; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; background-color: rgb(255, 255, 255); text-decoration-style: initial; text-decoration-color: initial; display: block;">

image

完整项目地址可以添加QQ源码群1004391443，有飞机大战、颜值打分器、打砖块小游戏、红包提醒神器、小姐姐表白神器等具体的实训项目，有清晰源码，有相应的文件

往下翻，发现这些都有注释，那就更好办了。

解析这个数据，拿出来hash值和filename，歌词lyric。

也没什么要说的了，直接贴代码

<pre spellcheck="false" style="box-sizing: border-box; margin: 5px 0px; padding: 5px 10px; border: 0px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-variant-numeric: inherit; font-variant-east-asian: inherit; font-weight: 400; font-stretch: inherit; font-size: 16px; line-height: inherit; font-family: inherit; vertical-align: baseline; cursor: text; counter-reset: list-1 0 list-2 0 list-3 0 list-4 0 list-5 0 list-6 0 list-7 0 list-8 0 list-9 0; background-color: rgb(240, 240, 240); border-radius: 3px; white-space: pre-wrap; color: rgb(34, 34, 34); letter-spacing: normal; orphans: 2; text-align: left; text-indent: 0px; text-transform: none; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; text-decoration-style: initial; text-decoration-color: initial;">import requests
from lxml import etree
import json
import re
import os
class kugou():
def startkugou(self):
for i in range(23, 24):
print(i)
res = requests.get('https://www.kugou.com/yy/rank/home/%s-8888.html?from=rank' % str(i))
self.get_song(res)
def get_song(self, res):
html = etree.HTML(res.content.decode('utf8'))
content = html.xpath('//script[10]')
content2 = content[0].text
# 解析出json列表，类型是str
content1 = content2.split('global.features =')[1].split('(function()')[0].strip()[0:-1]
try:
# 转换成json数据
content = json.loads(content1)
for i in range(len(content)):
hash = content[i]["Hash"]
file_name = content[i]["FileName"]
hash_url = "http://www.kugou.com/yy/index.php?r=play/getdata&hash=" + hash
hash_content = requests.get(hash_url)
play_url = ''.join(re.findall('"play_url":"(.?)"', hash_content.text))
lyrics = ''.join(re.findall('"lyrics":"(.?)"', hash_content.text))
real_download_url = play_url.replace("\", "")
try:
# if os.path.exists('kugou/' + file_name + '.txt'):
# print(file_name + " 歌词已经存在")
# # continue
# else:
with open('kugou/' + file_name + '.txt', 'w', encoding='utf8')as f:
f.write(lyrics.encode('utf8').decode('unicode_escape'))
print(file_name + "歌词已下载完成！")
# if os.path.exists('kugou/' + file_name + '.mp3'):
# print(file_name+" 歌曲已经存在")
# # continue
# else:
with open('kugou/' + file_name + ".mp3", "wb")as fp:
fp.write(requests.get(real_download_url).content)
print(file_name + "歌曲已下载完成！")
except OSError as e:
print("出现异常" + file_name)
file_name = self.validateTitle(file_name)
# if os.path.exists('kugou/' + file_name + '.txt'):
# print(file_name + " 歌词已经存在")
# # continue
# else:
with open('kugou/' + file_name + '.txt', 'w', encoding='utf8')as f:
f.write(lyrics.encode('utf8').decode('unicode_escape'))
print(file_name + "歌词已下载完成！")
# if os.path.exists('kugou/' + file_name + '.mp3'):
# print(file_name + " 歌曲已经存在")
# # continue
# else:
with open('kugou/' + file_name + ".mp3", "wb")as fp:
fp.write(requests.get(real_download_url).content)
print(file_name + "歌曲已下载完成！")
except json.decoder.JSONDecodeError as e:
print(e)
print(content2)
content1 = content2.split('global.features =')[1].strip().split('(function() {')[0].strip()
content1 = content1[0:-1]
print(content1)
def validateTitle(self, file_name):
""" 将 title 名字规则化
:param title: title name 字符串
:return: 文件命名支持的字符串 """
rstr = r"[=(),/\:*?"<>|' ']" # '= ( ) ， / \ : * ? " < > | ' 还有空格
new_title = re.sub(rstr, "_", file_name) # 替换为下划线
return new_title
if name == 'main':
kugou().startkugou()
</pre>

<tt-image data-tteditor-tag="tteditorTag" contenteditable="false" class="syl1556264619833 ql-align-center" data-render-status="finished" data-syl-blot="image" style="box-sizing: border-box; cursor: text; text-align: left; color: rgb(34, 34, 34); font-family: "PingFang SC", "Hiragino Sans GB", "Microsoft YaHei", "WenQuanYi Micro Hei", "Helvetica Neue", Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-ligatures: normal; font-variant-caps: normal; font-weight: 400; letter-spacing: normal; orphans: 2; text-indent: 0px; text-transform: none; white-space: pre-wrap; widows: 2; word-spacing: 0px; -webkit-text-stroke-width: 0px; background-color: rgb(255, 255, 255); text-decoration-style: initial; text-decoration-color: initial; display: block;">