Blogs

Querying the internet (100 billion rows) with MotherDuck

Motherduck

Common Crawl publishes petabytes of web crawl data on S3. With DuckDB and MotherDuck you can query the Common Crawl dataset directly, no cluster and no download, and measure how fast the vibe-coded web is growing.

Visit Site

Blogs Motherduck

BlogsWhy CSV Files Won’t Die and How DuckDB Conquers ThemMotherduck BlogsThe Future of BI: Exploring the Impact of BI-as-Code Tools with DuckDBMotherduck BlogsThis Month in the DuckDB Ecosystem: July 2025Motherduck BlogsSpecs Over Vibes: Consistent AI Results ft. Mark FreemanMotherduck BlogsWhat’s New: Streamlined User Management, Metadata, and UI EnhancementsMotherduck BlogsFuture Casting the Modern Data StackMotherduck Products & ServicesVerify Device Health with Security Posture Checks on 1Password1password BlogsMulti-Billion-Parameter Model Training Made Easy with CSoft R1.3 - CerebrasCerebras NewsCerebras Enables Notion to Deliver Real-Time Enterprise Search for 100+ Million Workspace UsersCerebras EventsData Outpost 2026Motherduck NewsCerebras Triples its Industry-Leading Inference Performance, Setting New All Time Record - CerebrasCerebras BlogsAWS PrivateLink for Neon Databases - NeonNeon BlogsWhat I Wish Someone Told Me When I Was Getting Into ARIASmashingmagazine ResearchNo News is Good News: A Critique of the One Billion Word BenchmarkCohere BlogsTesting if "bash is all you need"Vercel BlogsPreventing infrastructure abuse with Vercel FirewallVercel LearnBuilding An Offline-Friendly Image Upload SystemSmashingmagazine BlogsHow to send an email when updates are made to Google Sheets rowsZapier BlogsCanvas Is Now GA: AI-Guided Observability for Modern TeamsHoneycomb BlogsThe Internet of FunReplit