<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>博客 on Proooxy — 面向 AI Agent 与 LLM 的 RAG 就绪网页数据</title><link>https://proooxy.com/zh/blog/</link><description>Recent content in 博客 on Proooxy — 面向 AI Agent 与 LLM 的 RAG 就绪网页数据</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Wed, 01 Apr 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://proooxy.com/zh/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>2026 年网页抓取最佳实践：一线工程师指南</title><link>https://proooxy.com/zh/blog/web-scraping-best-practices-2026/</link><pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate><guid>https://proooxy.com/zh/blog/web-scraping-best-practices-2026/</guid><description>&lt;p&gt;在构建并维护 15 款生产级爬虫、服务超过 3,100 位用户、成功率保持 &amp;gt;99% 之后，以下是真正重要的实践经验。&lt;/p&gt;
&lt;h2 id="架构把爬虫当管道来设计而不是脚本"&gt;架构：把爬虫当管道来设计，而不是脚本&lt;/h2&gt;
&lt;p&gt;我见过最大的错误，就是把爬虫当成单步骤流程。生产级爬虫其实是数据管道：&lt;/p&gt;</description></item><item><title>读懂反爬虫防护：2026 年哪些方法依然有效</title><link>https://proooxy.com/zh/blog/bypassing-anti-bot-protection-guide/</link><pubDate>Sun, 15 Mar 2026 00:00:00 +0000</pubDate><guid>https://proooxy.com/zh/blog/bypassing-anti-bot-protection-guide/</guid><description>&lt;p&gt;反爬虫防护是一场军备竞赛。作为每天都要构建生产级爬虫来绕过这些系统的人，这里是我的一线实战视角——这些防护系统究竟在检查什么，合规的绕过技术又是什么样子。&lt;/p&gt;</description></item></channel></rss>