This past year and increasingly recently independent bookstores are being targeted by automated bot traffic on their sites. When this happens it causes error messages and slow loading web pages. We hear you, we share your deep frustration, and we continue to take aggressive steps to protect your digital storefronts.
What is "Web Scraping"?
Web scraping is an automated process where software programs (commonly called "bots" or "scrapers") are coded to visit websites at lightning speed to copy, extract, and steal data.
Your website holds a treasure trove of highly organized, clean, and comprehensive data—such as book titles, author names, ISBNs, pricing, inventory levels, and written descriptions. Bots are relentlessly pursuing this book data for a few key reasons:
- AI Training Data: Artificial Intelligence companies are aggressively harvesting massive datasets to train language models on structured, high-quality literary and catalog data.
- Aggressive Price Undercutting: Online marketplaces and third-party aggregators scrape independent bookstore data to run real-time market analysis, allowing them to automatically price-drop and undercut independent sellers.
- Amazon AI Shopping Agents: allows shoppers to purchase off-Platform goods directly through the Amazon app, sparking outrage among small and independent retailers who have had their inventory scraped by Amazon without permission. We’ve brought this to the attention of the FTC/DOJ (see Buy For Me search results on Google)
- New Start-ups: Attracted to our industry because of the recent growth of indie bookstores, are “innocently” using data from IndieCommerce bookstores as part of their model. This might include selling scraped book data for a monthly fee.
- Content Pirates: Specialized bots steal your rich book descriptions and book metadata to populate rival catalog systems or spam affiliate marketing sites.
- Bad Actors: These folks just want to cause problems for independent bookstores.
The Evolution of the Threat: Why Old Defenses are Failing
Historically, blocking bad bots was straightforward: they came from easily identifiable server centers, allowing legacy firewall systems to block them instantly.
Today, the landscape has shifted into an AI-driven crisis. In recent weeks, to avoid being blocked, these scrapers have taken on human-like browsing characteristics. They use thousands of residential IP addresses (routing their traffic through standard home Wi-Fi networks) and intentionally slow their speeds down to mimic how a real human browses your site. To a traditional security filter, a bot now looks exactly like a local customer clicking through your pages.
This is a Global, Ecommerce Battle
It is important to emphasize that this issue is not unique to our platform or our infrastructure. The largest tech giants in the world are currently fighting this exact same battle.
In fact, mainstream media and tech publications have recently reported heavily on a major legal showdown between Amazon and the AI search firm Perplexity. Amazon filed a federal lawsuit because Perplexity deployed specialized AI software agents to bypass Amazon's security, scraping product listings, pricing history, and customer reviews to feed their own databases.
Industry experts tracking this global shift have noted how difficult these human-mimicking bots are to stop:
"Since the traffic originates from a legitimate residential IP address, it is almost impossible for a website to distinguish a scraper's request from that of a genuine human user."
— AIMultiple, "The Most Common Web Scraping Challenges"
"Modern anti-bot systems have become very good at spotting automated patterns, which has pushed teams toward making the browser look as close to a real human session as possible at every signal point."
— 47Billion / Browserless, "State of Web Scraping"
Our Defense Strategy
Because these bots adapt continuously, we cannot rely on a single, static defense line. Over the past year alone, we have changed our blocking strategies over a dozen times. Each time we deploy a new blocking technique, the bot operators adjust their code to try and sneak past.
We are fighting back with a dynamic, multi-layered approach:
- Immediate Action: New Tech Beta Launching
We are currently working on an advanced bot-blocking technology with our hosting service provider and top CDN provider. This system looks beyond just IP addresses and evaluates complex browser "fingerprints" and microscopic behavioral anomalies. We began beta testing this new technology on our network on July 9th. We are seeing good preliminary results.
- Continuous Monitoring & Pattern Analysis
Our team is actively tracking scraping patterns around the clock, deploying rapid, live adjustments to intercept bots as their behavior evolves.
- You can alert us 24/7
If you or your customers are experiencing multiple error messages on your IndieCommerce website, or your pages are loading so slowly that your customer are unable to shop and your team can’t work on your website, please submit a Critical Outage Report
- Legislation & Advocacy
The ABA’s advocacy team is working on this as well by exploring potential legislation and bringing these incidents to the attention of the FTC and DOJ.
IndieCommerce is fully committed to protecting your data, your website's ability to serve your customers, and the integrity of your independent business. We will keep you closely updated on the results of our upcoming beta test. Thank you for your continued patience, vigilance, and partnership as we navigate this new digital frontier.
