作者busystudent (busystudent)
看板Python
标题[问题] 爬虫
时间Sun Nov 13 12:18:53 2016
各位好,想请问一段爬虫网址,一时之间不知道错在哪里,请大家指点一下,谢谢
问题点是我的div',{'class':'Titleinner'})#新闻标题和网址,怎麽都找不道资源了,之前撰写时可以挖出网址,
我退而检查SOUP,结果连个 div class Titleinner都没有挖出来,还请大家帮忙指点一下
好读版
https://ideone.com/BtvcXV
import requests
from bs4 import BeautifulSoup
name_list = ['Joshtwery']
for a in name_list:
links = ['
http://www.diigo.com/user/Joshtwery?page_num=0&type=all&sort=updated']
for link in links:
print link
res = requests.get(link,"lxml")
soup = BeautifulSoup(res.text.encode("utf-8"))
fol_table = soup.findAll('div',{'class':'Titleinner'})#新闻标题和网址
print fol_table#检查点
for a_link in fol_table:
a_links =str([tag['href']for tag in a_link.findAll('a',{'href':True})])[1:-1]
a_links =([tag['href']for tag in a_link.findAll('a',{'href':True})])
--
Sent from my Windows
--
※ 发信站: 批踢踢实业坊(ptt.cc), 来自: 1.172.114.184
※ 文章网址: https://webptt.com/cn.aspx?n=bbs/Python/M.1479010736.A.899.html
※ 编辑: busystudent (1.172.114.184), 11/13/2016 12:22:57
※ 编辑: busystudent (223.139.128.15), 11/13/2016 12:27:16
※ 编辑: busystudent (1.172.114.184), 11/13/2016 14:18:39
※ 编辑: busystudent (1.172.114.184), 11/13/2016 14:22:59
※ 编辑: busystudent (1.172.114.184), 11/13/2016 14:47:14
※ 编辑: busystudent (1.172.114.184), 11/13/2016 14:49:23
1F:→ s860134: 浏览器 F12 打开 本来你抓的页面就没有你要的东西 11/13 16:11
2F:→ s860134: 网页打开的内容是你爬的这个网址执行JS再去要的 11/13 16:11
3F:→ s860134: 自然爬这个只会爬到一陀程式码,里面甚麽东西都没有 11/13 16:12
4F:→ busystudent: 原来是这样 之前是没有js的 11/13 16:15
5F:→ busystudent: 那要怎麽处理js呢? 11/13 16:16
6F:→ s860134: 能直接抓 api 网址就直接抓,这个网站 API 很简单 11/13 16:29
7F:→ s860134: 浏览器 F12 录一下就有了,而且还 json 格式 11/13 16:29
8F:→ busystudent: wow说慢点 js我还是第一次处理 11/13 16:38
9F:→ busystudent: 我有找到js把原程式码藏在哪里 11/13 16:40
10F:→ busystudent: 就是你说录的那个地方 11/13 16:40
11F:→ busystudent: 我还在想下一步要怎麽做 11/13 16:41
12F:推 koshi0413: 直接当字串处理?小弟爬网页有Bs4,纯字串,axjx用开视 11/13 17:18
13F:→ koshi0413: 窗 11/13 17:18
14F:→ busystudent: 你好axjx怎麽开视窗呢? 11/13 17:19
15F:→ busystudent: 我之前的处理方式都是直接爬的 11/13 17:20
16F:推 koshi0413: 浏览器录网页,一个一个试,看是不是字串页面,不然就 11/13 17:27
17F:→ koshi0413: 用BS4,Bs4解不出来就去看用法,曾经为了一串标题用Bs4 11/13 17:27
18F:→ koshi0413: 半天才解掉 11/13 17:27
20F:推 koshi0413: 以s大的图片看,字串处理就可以了,re比较快 11/13 20:30
21F:→ busystudent: 我还是不懂,s大的东西怎麽来的,我网页录了一下,结果 11/13 20:59
23F:→ busystudent: 我之後该怎麽转成字串呢? 11/13 20:59
24F:→ busystudent: 另外s大的api网址是怎麽找到的 11/13 20:59
25F:→ s860134: 浏览器开发工具是你开启他後才开始录,你就打开後再重整 11/13 21:33
27F:→ s860134: josn 是网路通用格式,python 有内建 lib 去 parsing 11/13 21:38
28F:推 koshi0413: b大,照s大的方法解完後,花时间去网路找大数据学堂从 11/13 21:49
29F:→ koshi0413: 头看一下,会有帮助 11/13 21:49
30F:→ busystudent: 感谢楼上的各位大哥们! 11/13 22:32
31F:→ busystudent: 再请教叫一个问题, s大给的图片中第一行有一个api的 11/13 23:28
32F:→ busystudent: 网址,我一直都找不到是怎麽出来的 11/13 23:28
33F:→ busystudent: 我不管怎麽试都只有像是我第一张贴的图片那样 11/13 23:29
34F:→ hoho8: 怡?一下就找到啦。你切换到XHR试试,比较快找到 11/14 04:54