1. 携程景点评论抓取到底难在哪
先说清楚这篇要解决的事:用 Python 的 requests 加 BeautifulSoup,把携程景点详情页里的评分和评论抓下来,处理分页、请求头、返回数据格式不规整这些坑,最后把抓取结果做字段校验和限速验证。适合已经会写基础爬虫、但一碰到携程这种返回结构就卡住的人。
携程景点评论和普通静态网页不一样。你打开手机版携程搜一个景点,评论是滚动加载的,翻页靠的是 POST 请求往接口塞 JSON,返回的也不是干净的 HTML,而是一坨带转义、带\n、带\"的字符串。BeautifulSoup 在这里的作用不是解析标签,而是先把返回体当 HTML 过一遍,把转义字符和奇怪结构洗掉,再交给 json.loads。这个思路我第一次见也觉得别扭,但实测确实能跑通。
另一个坑是评论类型。综合评价、好评、差评是三个不同的入口,靠 payload 里CommentTagId区分:0 是综合,-11 是好评,-12 是差评。很多人只抓了综合就以为拿全了,其实好评差评要单独发请求。还有分页上限的问题,抓到 3000 条左右就翻不动了,这个后面单独讲。
整篇文章的脚本骨架你可以直接复制,请求头和 payload 都给你留好位置,改景点 ID 就能用。统一 Key 的部分用 TaoToken 管理,避免把各种 Key 散落在代码里。
2. 前置准备:TaoToken 统一 Key 与依赖安装
2.1 为什么要把 Key 收拢到一处
爬虫脚本里经常要调模型做评论情感分析、关键词抽取,或者让模型帮你判断返回的 JSON 结构对不对。如果每个脚本里都硬编码一个 Key,改起来就是灾难。TaoToken 的做法是给你一个统一的入口,模型对话、Coding Plan、API Keys 都在一个控制台里管。
官网入口:https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content=
API 地址是 https://taotoken.net/api ,注意这个不带 UTM 参数,配置里填这个就行。
2.2 config.toml 里放统一 Key
我习惯用 TOML 存配置,读起来清楚。在项目根目录建一个config.toml:
[taotoken] # 统一 Key,从控制台 API Keys 页面获取 api_key = "sk-你的统一Key" base_url = "https://taotoken.net/api" # 做评论分析时用的模型 model = "claude-sonnet" [spider] # 携程景点 BusinessId,换成你要抓的景点 business_id = "16588" # 起始页和结束页 page_start = 1 page_end = 5 # 每次请求间隔秒数,限速用 delay = 2.5读取配置的代码:
import tomllib with open("config.toml", "rb") as f: cfg = tomllib.load(f) API_KEY = cfg["taotoken"]["api_key"] BASE_URL = cfg["taotoken"]["base_url"] BIZ_ID = cfg["spider"]["business_id"]Python 3.11 以上自带 tomllib,低版本用pip install tomli然后import tomli as tomllib。
2.3 装依赖
pip install requests beautifulsoup4 lxmllxml 比 html.parser 快,但携程这个场景里我们只是拿 BeautifulSoup 洗字符串,用哪个解析器差别不大,html.parser 就够。
3. 可复制配置:请求头、payload 与解析脚本骨架
3.1 请求头怎么设
携程手机版接口对 User-Agent 有要求,用桌面版 UA 有时会返回空。我实测下面这套头比较稳:
HEADERS = { "User-Agent": ( "Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) " "AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1" ), "Content-Type": "application/json", "Referer": "https://m.ctrip.com/", "Origin": "https://m.ctrip.com", }Content-Type必须是application/json,因为 body 是 JSON。Referer 和 Origin 加上能降低被拦的概率。
3.2 payload 结构
payload 分两大块:CommentResultInfoEntity是查询条件,head是协议头。核心字段:
| 字段 | 含义 | 常用值 |
|---|---|---|
| BusinessId | 景点 ID | 换成目标景点 |
| BusinessType | 业务类型 | 11 表示景点 |
| ChannelType | 渠道 | 7 |
| CommentTagId | 评论类型 | 0 综合 / -11 好评 / -12 差评 |
| PageIndex | 页码 | 从 1 开始 |
| PageSize | 每页条数 | 10 |
| SortType | 排序 | 3 |
构造 payload 的函数:
def build_payload(biz_id, page_index, tag_id=0, page_size=10): return { "CommentResultInfoEntity": { "BusinessId": str(biz_id), "BusinessType": 11, "ChannelType": 7, "CommentTagId": tag_id, "ImageFilter": "false", "PageIndex": page_index, "PageSize": page_size, "PoiId": 0, "SortType": 3, "StarType": 0, "TouristType": 0, "VideoImageHeight": 392, "VideoImageWidth": 700, }, "contentType": "json", "head": { "auth": "", "cid": "09031081211299101374", "ctok": "", "cver": "1.0", "extension": [{"name": "protocal", "value": "https"}], "lang": "01", "sid": "8888", "syscode": "09", }, }3.3 返回体清洗:为什么需要 deal_json_invalid
携程返回的 body 里,JSON 的引号被转义成了\",还有大量\n、\r、\t,甚至嵌套的"{"这种结构。直接 json.loads 会报错。原作者的思路是用一串 replace 把特殊结构先替换成占位符,再换回来。我把它整理成一个更可读的版本:
def clean_response(text): text = text.replace("\n", "").replace("\r", "").replace("\t", "") text = text.replace('\\"', '"') # 处理嵌套引号结构 text = text.replace('":"{"', '**P5**') text = text.replace('":"', '&&P&&') text = text.replace('","', '$$P$$') text = text.replace('":{"', '**P1**') text = text.replace('"},"', '**P2**') text = text.replace(',"', '**P3**') text = text.replace('{"', '@@P@@') text = text.replace('"}', '**P4**') text = text.replace('":', '**P6**') text = text.replace('"', '”') # 还原 text = text.replace('**P5**', '":{"') text = text.replace('&&P&&', '":"') text = text.replace('$$P$$', '","') text = text.replace('**P1**', '":{"') text = text.replace('**P2**', '"},"') text = text.replace('@@P@@', '{"') text = text.replace('**P4**', '"}') text = text.replace('**P3**', ',"') text = text.replace('**P6**', '":') text = text.replace('\\"', '"') text = text.replace('}"', '}') return text注意:这套 replace 是原作者针对当时接口格式写的,接口结构可能变。如果清洗后 json.loads 还是报错,先把原始 text 打印出来看具体哪里不对,再针对性加 replace,别盲目套。
3.4 完整抓取骨架
import json import time import datetime import requests from bs4 import BeautifulSoup API_URL = ( "https://m.ctrip.com/restapi/soa2/13444/json/" "GetCommentListAndHotTagList?_fxpcqlniredt=09031081211299101374" "&__gw_appid=99999999&__gw_ver=1.0&__gw_from=10650019636&__gw_platform=H5" ) def fetch_page(biz_id, page_index, tag_id=0): payload = build_payload(biz_id, page_index, tag_id) resp = requests.post(API_URL, json=payload, headers=HEADERS, timeout=15) resp.raise_for_status() soup = BeautifulSoup(resp.text, "html.parser") cleaned = clean_response(str(soup)) return json.loads(cleaned) def parse_comments(data): rows = [] comments = data.get("CommentResult", {}).get("CommentInfo", []) for c in comments: ts = int(c["PublishTime"][6:16]) dt = datetime.datetime.utcfromtimestamp(ts).strftime("%Y-%m-%d %H:%M:%S") user = c.get("UserInfoModel") or {} rows.append({ "total_star": c.get("TotalStar"), "publish_time": dt, "district": user.get("UserDistrictName", "None"), "medal": user.get("MedalName", "None"), "content": c.get("Content", ""), }) return rows def crawl(biz_id, page_start, page_end, tag_id=0, delay=2.5): all_rows = [] for page in range(page_start, page_end + 1): data = fetch_page(biz_id, page, tag_id) rows = parse_comments(data) if not rows: print(f"第 {page} 页无数据,停止") break all_rows.extend(rows) print(f"第 {page} 页抓到 {len(rows)} 条") time.sleep(delay) return all_rows4. 验证请求:跑通一次并校验字段
4.1 先跑单页
if __name__ == "__main__": rows = crawl(BIZ_ID, 1, 1, tag_id=0, delay=0) for r in rows[:3]: print(r)正常输出类似:
{'total_star': 5, 'publish_time': '2024-03-12 08:21:00', 'district': '上海', 'medal': '金牌会员', 'content': '景色很美,值得一去'}4.2 字段校验
抓完一批后,做几个断言,确认数据没缺:
def validate(rows): assert rows, "结果为空" for r in rows: assert r["total_star"] is not None, "评分缺失" assert r["publish_time"], "时间缺失" assert isinstance(r["content"], str), "评论内容类型错误" print(f"校验通过,共 {len(rows)} 条")4.3 限速验证
限速不是随便 sleep 就行,要验证 sleep 是否真的生效。记录每次请求的时间戳:
import time def crawl_with_timing(biz_id, page_start, page_end, delay): stamps = [] for page in range(page_start, page_end + 1): t0 = time.time() fetch_page(biz_id, page) stamps.append(time.time() - t0) time.sleep(delay) gaps = [stamps[i+1] - stamps[i] for i in range(len(stamps)-1)] print("请求间隔:", [round(g, 2) for g in gaps]) assert all(g >= delay * 0.9 for g in gaps), "限速未生效"如果间隔明显小于 delay,说明 sleep 位置放错了,或者请求本身太快导致时间戳计算有偏差。
5. 本篇常见错排查
5.1 json.loads 报 JSONDecodeError
最常见。原因基本是 clean_response 没洗干净。排查步骤:先把resp.text存到文件,用编辑器打开看结构,找到报错位置附近的字符,针对性加 replace。别一上来就怀疑接口挂了。
5.2 抓到 3000 条就翻不动了
原作者也提到这个现象。携程对分页深度有限制,PageIndex 超过某个值后返回空列表或重复数据。应对方式:一是按时间分段抓,用不同 SortType 或时间范围缩小每次查询;二是好评差评分开抓,每个类型单独翻页,总量能多拿一些。别指望一个 tag 翻到底。
5.3 返回空 CommentInfo
检查 BusinessId 和 BusinessType 是否匹配。景点是 11,酒店是别的值,混用会返回空。另外 CommentTagId 填错也会空,0/-11/-12 三个值确认一下。
5.4 请求被拒或返回验证页
大概率是请求头不对或频率太高。把 User-Agent 换成手机版,加上 Referer 和 Origin,delay 调到 3 秒以上。如果还不行,检查是不是短时间内请求太密集,停一会儿再试。
5.5 PublishTime 解析出错
携程返回的是/Date(1234567890000)/这种格式,取中间数字部分转时间戳。如果切片位置不对,先打印原始字符串确认格式,再调整[6:16]的索引。
6. 把 Key 和抓取流程接起来
抓完评论后,通常要做情感分析或关键词提取。这时候用 TaoToken 的统一 Key 调模型,不用在爬虫脚本里再塞一个 Key。配置已经在 config.toml 里,直接读:
import requests def analyze_comment(text): resp = requests.post( f"{BASE_URL}/v1/messages", headers={ "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json", }, json={ "model": cfg["taotoken"]["model"], "messages": [{"role": "user", "content": f"判断这条评论情感倾向:{text}"}], }, timeout=30, ) return resp.json()模型对话入口在 https://taotoken.net/api-keys ,API Keys 页面能拿到统一 Key。如果你要长期跑编码任务或者 Agent 流程,Coding Plan 更适合,入口是 https://taotoken.net/coding-plan 。接入文档在 https://taotoken.net/doc ,里面有各语言的调用示例。
抓取脚本本身不依赖模型,但把评论清洗、字段校验、情感分析串成一条流水线时,统一 Key 能省掉很多配置切换的麻烦。我一般把爬虫结果先落成 JSONL,再批量送模型分析,这样爬和析解耦,哪一步出问题都好定位。