Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright
DEV Community

Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright

Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering. A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise. Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering. 1. The Bottleneck: Commercial APIs vs. Fragile Scrapers Most teams start with simple libraries like BeautifulSoup or newspaper3k . These quickly break down: - Layout Brittleness: Custom CSS selectors degrade when publishers redesign their DOM or randomize class names (e.g., CSS Modules/Tailwind compiles). - Client-Side Rendering (CSR): Single-page applications built on React, Next.js, or Vue return empty shells to standard HTTP clients. - Commercial SaaS Overkill: Hosted extraction APIs charge upwards of $0.002 to $0.01 per page. If your pipeline indexes 200,000 URLs monthly, you are paying hundreds of dollars for what essentially amounts to an HTTP request and an AST traversal. To balance latency, compute cost, and reliability, the optimal architecture uses a tiered extraction waterfall: - Tier 1 (Fast Path): High-speed async HTTP fetch + trafilatura . Latency: ~150-300ms. - Tier 2 (Heuristic Fallback): If Trafilatura fails or yields empty strings, fall back to Mozilla's readability-lxml algorithm. - Tier 3 (Headless Browser Fallback): If client-side hydration or dynamic script loading is detected, trigger an isolated headless Playwright worker to render the DOM before extraction. 2. System Architecture [Incoming Request: URL] │ โ–ผ ┌───────────────────┐ │ Redis Cache Check │ ──(Hit)──โ–บ Return Markdown & Metadata └───────────────────┘ │ (Miss) โ–ผ ┌───────────────────┐ │ HTTPX Async Fetch │ └───────────────────┘ │ [Static HTML] │ โ–ผ ┌───────────────────┐ │ Trafilatura │ ──(Success: Length > Threshold)──โ–บ Parse Meta & Return └───────────────────┘ │ (Failed / Empty) โ–ผ ┌───────────────────┐ │ Readability-lxml │ ──(Success: Length > Threshold)──โ–บ Parse Meta & Return └───────────────────┘ │ (Failed / SPA Detected) โ–ผ ┌───────────────────┐ │ Playwright Worker │ ──(Render DOM)──โ–บ Re-extract via Trafilatura └───────────────────┘ │ โ–ผ [Store in Cache & Return Markdown] This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero. 3. Implementation: Code & Core Logic Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution. 3.1 Pydantic Validation & Schemas from pydantic import BaseModel, HttpUrl, Field from typing import Optional, Dict, Any from datetime import datetime class ExtractionRequest(BaseModel): url: HttpUrl force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering") max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits") class PageMetadata(BaseModel): title: Optional[str] = None author: Optional[str] = None published_date: Optional[str] = None site_name: Optional[str] = None reading_time_minutes: int = 0 opengraph: Dict[str, Any] = {} class ExtractionResponse(BaseModel): url: str markdown: str metadata: PageMetadata engine_used: str execution_time_ms: float 3.2 The Dual-Engine Core Service import time import httpx import trafilatura from readability import Document from bs4 import BeautifulSoup from markdownify import markdownify as md from playwright.async_api import async_playwright USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36" async def fetch_html_fast(url: str) -> str: async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client: response = await client.get(url, headers={"User-Agent": USER_AGENT}) response.raise_for_status() return response.text async def fetch_html_playwright(url: str) -> str: async with async_playwright() as p: browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"]) context = await browser.new_context(user_agent=USER_AGENT) page = await context.new_page() await page.goto(url, wait_until="networkidle", timeout=20000) content = await page.content() await browser.close() return content def extract_metadata(html: str) -> PageMetadata: soup = BeautifulSoup(html, "lxml") og_data = {} for tag in soup.find_all("meta"): prop = tag.get("property", tag.get("name", "")) if prop.startswith("og:") or prop.startswith("twitter:"): og_data[prop] = tag.get("content", "") title = og_data.get("og:title") or (soup.title.string if soup.title else None) author = og_data.get("article:author") or og_data.get("twitter:creator") pub_date = og_data.get("article:published_time") return PageMetadata( title=title, author=author, published_date=pub_date, site_name=og_data.get("og:site_name"), opengraph=og_data ) def extract_content(html: str, url: str) -> tuple[str, str]: # Primary Engine: Trafilatura extracted = trafilatura.extract( html, url=url, output_format="markdown", include_links=True, include_images=False, favor_recall=False ) if extracted and len(extracted.strip()) > 200: return extracted, "trafilatura" # Fallback Engine: Readability + Markdownify doc = Document(html) summary_html = doc.summary() markdown_output = md(summary_html, heading_style="ATX").strip() if len(markdown_output) > 100: return markdown_output, "readability" return "", "none" 3.3 The FastAPI Microservice Route from fastapi import FastAPI, HTTPException, status app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0") @app.post("/api/v1/extract", response_model=ExtractionResponse) async def extract_article(payload: ExtractionRequest): start_time = time.perf_counter() url_str = str(payload.url) engine_used = "trafilatura" try: if payload.force_playwright: html = await fetch_html_playwright(url_str) markdown, engine_used = extract_content(html, url_str) engine_used = f"playwright+{engine_used}" else: # Tier 1 & 2: Fast HTTP path try: html = await fetch_html_fast(url_str) markdown, engine_used = extract_content(html, url_str) except Exception: markdown = "" # Tier 3: SPA / Fallback if empty if not markdown or len(markdown.strip()) payload.max_chars: markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]" exec_duration = (time.perf_counter() - start_time) * 1000 return ExtractionResponse( url=url_str, markdown=markdown, metadata=metadata, engine_used=engine_used, execution_time_ms=round(exec_duration, 2) ) except HTTPException: raise except Exception as e: raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e)) 4. Deployment, Caching & Concurrency Hardening When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards: Browser Resource Constraints Launching a new Chromium instance via async_playwright on every request will instantly lead to Linux OOM crashes. In a production build: - Maintain a single long-lived Browser instance. - Create and destroy lightweight BrowserContext sessions per request. - Enforce --disable-dev-shm-usage and--no-sandbox flags inside your Docker container. Redis Caching Layer Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests: import hashlib def generate_cache_key(url: str) -> str: normalized = url.split("?")[0].rstrip("/").lower() return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}" Docker Container Configuration FROM python:3.11-slim ENV PYTHONUNBUFFERED=1 \ PIP_NO_CACHE_DIR=1 \ PLAYWRIGHT_BROWSERS_PATH=/ms-playwright WORKDIR /app RUN apt-get update && apt-get install -y --no-install-recommends \ libxml2-dev libxslt-dev gcc libc-dev curl \ && rm -rf /var/lib/apt/lists/* COPY requirements.txt . RUN pip install -r requirements.txt RUN playwright install --with-deps chromium COPY . . EXPOSE 8000 CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"] 5. Conclusion & Ready-to-Use Workflow You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above. If you want the complete, production-ready implementation out of the box-including: - Pre-configured Docker Compose with Redis cache and persistent volume configurations - Playwright connection-pooling to avoid concurrency memory leaks - Rate-limiting middleware (slowapi) and health check probes - Complete test suites with synthetic fixtures and ready-to-import n8n orchestration workflows You can grab the complete source repository directly: - Instant Access on Whop: Get the Microservice on Whop - Direct Download on Gumroad: Download on Gumroad - Use promo code EARLYBIRD at checkout for 20% off. Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.