r/Rag
Web Scraping for RAG System: How to Handle Client-Side Rendered Sites
- upvotes
- 6
- comments
- 13
Post
The quality of RAG pipeline is mainly bounded by the cleanliness of its embeddings and if you feed poorly formatted data into vector database, retrieval accuracy drops drastically. The problem is that most of the modern docs and web sources are increasingly client side rendered (CSR) and when you scrape them for RAG pipelines, it usually run into two major issues: Standard scraping tools pull before JS finishes running and leaving you with empty <div id="__next"></div> chunks If you embed raw HTML, 60% of your vector similarity is matching against things like CSS classes, navigation links and tracking scripts rather than actual technical answers So after spending a lot of time behind doing this, I wanted to share a few architecture to scrape client rendered sites for RAG: - Wait for client hydration before extraction and use a scraper with DOM stability checks and not static timeouts and if a page loads dynamically, your scraper must wait until background fetch calls resolve and the DOM tree stabilizes - Convert directly to semantic markdown before chunking cause its optimal format for LLM embeddings headers (#, ##), bullet points and code blocks preserve semantic hierarchy without adding token overhead so strip all inline styles, SVGs and header/footer navigation - Offload ingestion to Firecrawl instead of building a pipeline of playwright -> readability -> turndown...
Keep reading with a free account
The rest of this post, and every signal for LangChain, is in your free account.
Extracted from these lines
Chunk by markdown headers and not arbitrary character counts so don't use a fixed 500 character splitter that cuts sentences or code blocks in half but you can use the Split on markdown headers (MarkdownHeaderTextSplitter in langchain/llamaIndex) so each chunk represents a cohesive section
From the post
Offload ingestion to Firecrawl instead of building a pipeline of playwright -> readability -> turndown -> langchain, firecrawl’s /crawl or /scrape endpoint are great...
From the post