So, I was poking around figma.com and noticed they’ve got this line in their robots.txt:
User-Agent: Google-Extended
Disallow: /
Basically, this tells Google to stay out of their content, especially for Gemini AI training. It's a smart move, right? But I had to wonder, is it really worth it?
What I found out is that if you want to check your own site, just go to yourdomain.com/robots.txt, hit Ctrl-F, and search for “Google-Extended.” It takes like ten seconds. Seriously, it’s that easy.
Most folks who added that line did so to keep their content from being used by AI. Fair play to them. But here’s the kicker: Google claims blocking it doesn’t impact your site’s performance in Search or AI Overviews. All it does is control how Gemini gets trained.
But there’s a study indicating that sites blocking it were pulled less often by both Gemini and those AI Overviews compared to sites that didn’t block it. Now, before you jump to conclusions, correlation doesn’t equal causation. Sometimes, sites blocking one crawler are just more locked down overall, and maybe that’s why they get less love.
What’s evident is that businesses who rely on being discovered are doing the opposite. Take Zapier; they’ve embraced AI and even added signals like ai-train=yes, search=yes, ai-input=yes. They’re inviting Google’s AI in with open arms.
The crux of it all? You’ve got to decide what you want. If you want AI to recommend your content, you might consider letting it in and giving it something juicy to quote—like pages that really tackle your buyers' questions with some solid comparisons and data. But if protecting your content is your priority, then keep that block in place. Just remember, don't let it be a random copy-paste job from 2023.
Once you allow AI access, you better have something worth citing. That’s the whole idea behind Outrank: we create answer-shaped articles designed to go straight to your CMS. Give it a spin for free.
