seen from United States

seen from China

seen from Germany

seen from Türkiye
seen from China

seen from Malaysia
seen from Philippines

seen from United States
seen from Saudi Arabia
seen from Russia
seen from China
seen from China

seen from Netherlands

seen from Japan

seen from United States
seen from United States
seen from Australia
seen from France
seen from Malaysia
seen from Singapore
bot-ched
Okay, so I'm the Growth Team product manager at my job - so my goal is getting more people to see the site. In this day and age, that means SEO: search engine optimization.
Even with the scales removed (sorry - confidential info!) it's clear things are going up. Well... until we hit a snag. See, Google doesn't like duplicate content, but Googlebot - the webcrawler that visits your site and reports back - is pretty dumb. It thinks example.com and example.com?session_id=123 are two different sites, even though a human could see this query parameter has nothing to do with the contents of the site. So... we gotta give Googlebot some rules.
Now, I've never bothered to do this for my own projects, but it's best practice to include a top-level robots.txt file so those non-human visitors go where they're supposed to. For example, we recently started disallowing query parameters to eliminate the duplicate page issue.
Uh... as it turns out, due to our site’s setup, one particular query parameter is exceptionally important, specifically for Googlebot. In fact, if the robot visits the page without that, it sees an empty shell. And our search traffic declines, week-over-week:
Luckily, the team figured it out and fixed it right away - in a single line of code! - and this past week had some of our highest-traffic days ever. But let this be a lesson: take extreme care when it comes to the bots!
tl;dr: duplicate content is bad, but no content is worse; use robots.txt, and give the bots a sitemap too; track it all in Google Analytics
project: this is my job, actually
problem patterns
It's been a busy couple weeks at my job, where I'm now working across three separate engineering teams (and all the projects that entails). I've also been coding a lot, and we'll get back to that, but first I thought I'd share something interesting from the Growth Team.
Google Search Console is a tool that lets you see what Googlebot sees as it tries to index your website. In our case, it saw a not-insignificant number of 404s - that is, links on our pages directing traffic to non-pages. Google's guidelines say this ain't so bad, but I bet it becomes a problem when we're getting as many as we’ve got.
So... I need to figure out why. And I certainly saw some problematic patterns in our processes - for example, when an old article is deleted, it's up to an individual (non-technical) writer to create a "redirect" to a new one - or a listing page - to avoid 404s. Stuff like that. But that's not what's happening here:
See, our custom content management system (ie, the analog of the editor I'm typing this Tumblr post on) interprets HTML, meaning links look like this:
See the problem? Since there's no automatic link formatting or validating, writers can accidentally create internal links (to non-existent pages) when they mean to make external links (to other websites). So I spent a few hours one morning coming up with a regex pattern to catch these guys:
Obviously, this catches the problem after the fact, when the links show up in HTML, but our engineers were able to make it more useful by making sure the markdown includes a protocol/scheme (http://, https://, mailto:, etc.). The fix isn't in yet, but fingers crossed, our 404s go missing.
tl;dr: treat the disease, not the symptoms; Google Search Console is a great error-catching tool; regular expressions come in handy
project: our 404 page, I guess?
How do search engines work? Learn the step-by-step technical process of how search engines crawl, render, index, and rank websites to delive
How Search Engines Crawl and Index Your Website for SEO
Understanding the Crawling Process
Search engines use automated bots, often called spiders or crawlers, to navigate the internet and discover content. These bots follow links from known pages to new ones, systematically mapping the web to find both new and updated information.
Crawlers like Googlebot start with a list of URLs from previous crawls.
They follow links on these pages to discover new locations.
Your site structure and internal linking help crawlers move efficiently through your content.
How Indexing Organizes Your Data
Once a search engine crawls your site, it processes the information to store it in a massive database known as the index. During this phase, the engine analyzes the text, images, and video files to understand the context and subject matter of each page.
Indexing is essentially the search engine cataloging your site.
Content must be indexed to appear in any search results page.
Using a sitemap helps search engines find your important pages faster.
Improving Your Search Visibility
If a search engine cannot crawl or index your site, your content remains invisible to users regardless of its quality. Optimizing for technical SEO ensures that search bots have easy access to your pages so they can be processed and ranked.
Ensure your robots.txt file does not block critical content.
Fix broken links to help crawlers navigate your site without dead ends.
Focus on high-quality content that provides value once it is indexed.
Cloudflare Inc News in Quantum Security Shape 2025 Internet
Cloudflare predicts 2025 internet trends: Rising Cyberwarfare, Bot Wars, Quantum Security News from Cloudflare
Cloudflare Inc., the leading connectivity cloud service, presented its sixth annual Year in Review today, providing one of the most comprehensive examinations of traffic, security, and global Internet trends in 2025. The research, based on data from Cloudflare's worldwide network in over 330 cities in 120 countries, highlights a pivotal year online with rapid technological progress and rising risks.
Society relies on the Internet for personal and professional reasons, thanks to technological advances. Worldwide Internet traffic rose 19% last year. Cloudflare CEO Matthew Prince says artificial intelligence and more creative threat actors are “fundamentally rewired” the Internet. Emergence of AI and Bot Wars As AI bot fights intensified this year, Google's crawling bot ruled. Googlebot, which crawls for search indexing and AI training, generated the most automated Internet traffic. More than 28% of Verified Bot traffic in 2025 comes from Googlebot. GoogleBot contributed 4.5% of HTML request traffic, somewhat more than the 4.2% of all AI bots. ChatGPT/OpenAI led Generative AI's rapid growth. Google Gemini, Grok/xAI, and DeepSeek joined the top 10 list, while Perplexity, Claude/Anthropic, and GitHub Copilot rose. From an AI crawling perspective, model training accounts for most traffic. By 2025, user action crawling, in which bots examine websites in response to user enquiries to chatbots, had the lowest volume but the greatest growth, expanding more than 15 times. Due to this surge, site owners often added completely forbidden directives to robots.txt files, making AI crawlers the most restricted user agents. On Cloudflare's Workers AI development platform, text production was the most common activity for Meta's llama-3-8b-instruct model. Security Achievements During Cyber Escalation More than 25 record-breaking DDoS attacks arose from the rise in cyberwarfare caused by global traffic. This growth in volumetric attacks redefined online danger "scale". The most targeted vertical was “People and Society”—civil associations, non-profits, and religious institutions—with 4.4% of global reduced traffic. This industry was the first to be attacked, possibly due of its sensitive user data and potential financial value. Post-quantum encryption, which secures 52% of human traffic, arrived quickly, marking a security milestone. Innovative quantum computing poses risks to customers, hence this strategy is necessary. Post-quantum encrypted traffic jumped from 29% to 52% worldwide this year. Apple's mid-September operating system modifications allowed TLS-protected connections to automatically advertise TLS 1.3 quantum-secure key exchange compatibility, accelerating its adoption. Global Quality and Connectivity Leaders Governments caused over half of the 174 significant Internet outages worldwide in 2025. Regional and national shutdowns were used to dissuade academic exam cheating in Iraq, Syria, and Sudan. However, cable-cut outages fell roughly 50%. Europe has the best connectivity, with download rates topping 200 Mbps. Spain had the best Internet quality worldwide with download speeds above 300 Mbps and upload speeds up to 206 Mbps. The UNICO-Broadband initiative, which aspires to build infrastructure with symmetric speeds of at least 300 Mbps, may explain this amazing performance. Starlink's satellite Internet service received more requests from over 20 new countries and regions this year. Mobile requests rose to 43% worldwide, with over half of all request traffic coming from mobile devices in 117 countries and regions. When considering the technology stack of modern web projects, Go-based clients made 20% of automated API calls, up from 12% in 2024, followed by Python at 17%.
Robots.txt
El Portero de tu Web que Debes Conocer 🤖 ¡Hola a todos los entusiastas de la tecnología en Alicante y más allá! Hoy en n’tics vamos a hablar de algo que suena a película de ciencia ficción pero que es fundamental para tu página web: el archivo robots.txt. ¿Un archivo que da órdenes a los robots? ¡Pues sí, algo así! Y es más fácil de entender de lo que crees. ¡Vamos al lío! ¿Qué es el archivo…
Google Confirms: No LLMS.txt Needed for AI Overviews
Google’s Gary Illyes says LLMS.txt is not used for ranking in AI Overviews. Instead, follow standard SEO practices to appear in AI results. Google Says LLMS.txt Doesn’t Matter for AI Rankings — Stick to Normal SEO, Says Gary Illyes In a clear message to content creators and SEO professionals, Google has confirmed that normal SEO practices are all you need to appear in AI Overviews — its…
How to Fix Crawl Budget Waste for Large E-Commerce Sites
Struggling with crawl budget waste on your massive e-commerce site?
Learn actionable strategies to fix crawl budget waste for large e-commerce sites, optimize Googlebot’s efficiency, and boost your SEO rankings without breaking a sweat.
Introduction: When Googlebot Goes on a Wild Goose Chase 🕵️♂️
Picture this: Googlebot is like an overworked librarian trying to organize a chaotic library. Instead of shelving bestsellers, it’s stuck rearranging pamphlets from 2012.
That’s essentially what happens when your e-commerce site suffers from crawl budget waste.
Your precious crawl budget—the number of pages Googlebot can and will crawl on your site—gets squandered on irrelevant, duplicate, or low-value pages. Yikes!
For large e-commerce platforms with millions of URLs, this isn’t just a minor hiccup; it’s a full-blown crisis.
Every second Googlebot spends crawling a broken filter page or a duplicate product URL is a second not spent indexing your shiny new collection.
So, how do you fix crawl budget waste for large e-commerce sites before your SEO rankings take a nosedive? Buckle up, buttercup—we’re diving in.
What the Heck Is Crawl Budget, Anyway? (And Why Should You Care?) 🤔
H2: Understanding Crawl Budget: The Lifeline of Your E-Commerce SEO
Before we fix crawl budget waste for large e-commerce sites, let’s break down the basics. Crawl budget refers to the number of pages Googlebot will crawl on your site during a given period. It’s determined by:
Crawl capacity limit: How much server strain Googlebot is allowed to cause.
Crawl demand: How “important” Google deems your site (spoiler: high authority = more crawls).
For e-commerce giants, a limited crawl budget means Googlebot might skip critical pages if it’s too busy crawling junk. Think of it like sending a scout into a maze—if they waste time on dead ends, they’ll never reach the treasure.
How to Fix Crawl Budget Waste for Large E-Commerce Sites: 7 Battle-Tested Tactics
1. Audit Like a Bloodhound: Find What’s Draining Your Budget 🕵️♀️
First things first—you can’t fix what you don’t understand. Run a site audit to uncover:
Orphaned pages: Pages with no internal links. (Googlebot can’t teleport, folks!)
Thin content: Product pages with 50-word descriptions. Cue sad trombone.
Duplicate URLs: Color variants? Session IDs? Parameter hell? Fix. Them.
Broken links: 404s and 500s that send Googlebot into a loop.
Pro Tip: Use Screaming Frog or Sitebulb to crawl your site like Googlebot. Export URLs with low traffic, high bounce rates, or zero conversions. These are prime suspects for crawl budget waste.
2. Wield the Robots.txt Sword (But Don’t Stab Yourself) ⚔️
Blocking Googlebot from crawling useless pages is a no-brainer. But tread carefully—misconfigured robots.txt files can backfire. Here’s how to do it right:
Block low-priority pages: Admin panels, infinite pagination (page=1, page=2…), and internal search results.
Avoid wildcard overkill: Disallow: /*?* might block critical pages with parameters.
Test with Google Search Console: Use the robots.txt tester to avoid accidental blockages.
3. Canonical Tags: Your Secret Weapon Against Duplicates 🔫
Duplicate content is the arch-nemesis of crawl budget. Fix it by:
Adding canonical tags to all product variants (e.g., rel="canonical" pointing to the main product URL).
Using 301 redirects for deprecated or merged products.
Consolidating pagination with rel="prev" and rel="next" (though Google’s support is spotty—proceed with caution).
4. XML Sitemaps: Roll Out the Red Carpet for Googlebot 🎟️
Your XML sitemap is Googlebot’s GPS. Keep it updated with:
High-priority pages: New products, seasonal collections, bestsellers.
Exclude junk: No one needs 50 versions of the same hoodie in the sitemap.
Split sitemaps: For sites with 50k+ URLs, split into multiple sitemaps (e.g., products, categories, blogs).
5. Fix Internal Linking: Turn Your Site into a Well-Oiled Machine ⚙️
A messy internal linking structure forces Googlebot to play hopscotch. Optimize by:
Adding breadcrumb navigation for layered category pages.
Linking to top-performing pages from high-authority hubs (homepage, blogs).
Pruning links to low-value pages (looking at you, outdated promo codes).
6. Dynamic Rendering: Trick Googlebot into Loving JavaScript 🎭
Got a JS-heavy site? Googlebot might struggle to render pages, leading to crawl inefficiencies. Dynamic rendering serves a static HTML snapshot to bots while users get the full JS experience. Tools like Prerender or Puppeteer can help.
7. Monitor, Tweak, Repeat: Crawl Budget Optimization Is a Marathon 🏃♂️
Fixing crawl budget waste isn’t a one-and-done deal. Use Google Search Console to:
Track crawl stats (pages crawled/day, response codes).
Identify sudden spikes in 404s or server errors.
Adjust your strategy quarterly based on data.
FAQs: Your Burning Questions, Answered 🔥
Q1: How often should I audit my site for crawl budget waste?
A: For large e-commerce sites, aim for quarterly audits. During peak seasons (Black Friday, holidays), check monthly—traffic surges can expose new issues.
Q2: Can crawl budget waste affect my rankings?
A: Absolutely! If Googlebot’s too busy crawling junk, your new pages might not index quickly, hurting visibility and sales.
Q3: Are pagination pages always bad?
A: Not always—but if they’re thin or duplicate, block them with robots.txt or consolidate with View-All pages.
Conclusion: Stop the Madness and Take Back Control 🛑
Fixing crawl budget waste for large e-commerce sites isn’t rocket science—it’s about playing smart with Googlebot’s time. By auditing ruthlessly, blocking junk, and guiding bots to your golden pages, you’ll transform your site from a chaotic maze into a well-organized powerhouse. Remember, every crawl Googlebot makes should count. So, roll up your sleeves, implement these tactics, and watch your SEO performance soar. 🚀
Still sweating over crawl budget issues? Drop a comment below—we’ll help you troubleshoot. Fix All Technical Issus Now