Data-Intensive SEO Platform
Since 2017, Batunet has been responsible for the technical development of a commercial, data-intensive SEO platform. We took over a product that was already live, with a much smaller feature set. Since then, Batunet has newly built core subsystems and continuously extended existing areas.
The case in one minute.
- Role
- Technical development of the platform since 2017 — concept, architecture, implementation and operations.
- Starting point
- A live, commercially used product with a much smaller feature set.
- Product type
- Commercial SaaS product.
- Data volume
- Billions of records, dataset in the terabyte range.
- Built entirely by Batunet
- Technical website crawler, reporting system, AI feature set.
- Processing
- Distributed and asynchronous processing via workers and queues, plus custom job scheduling.
- Technical environment
- PHP, Symfony, MySQL/MariaDB, Redis, RabbitMQ, Elasticsearch/OpenSearch, JavaScript, TypeScript, Docker, Linux/Nginx.
- Division of work
- The operator defines business requirements and product goals. The technical concept is largely developed by Batunet.
Evolve, don’t start over.
The product was running when we took it over. It was live, commercially used and had a working core — with a feature set much smaller than today’s. Starting over in that situation means throwing away proven domain logic only to build it again.
So we didn’t start over. Every extension went into the running system while it was in production. The core has stayed; the feature set has grown continuously since 2017: core subsystems were newly built, existing areas extended.
What was built here.
Three subsystems were built entirely by Batunet, another one almost entirely. Existing areas were taken over and significantly extended.
- Built entirely by Batunet
- Technical website crawler, Reporting system, AI feature set
- Built almost entirely by Batunet
- Export system
- Significantly extended
- Rank monitoring, Keyword analysis, Backlink analysis, Competitor analysis, Content features, Monitoring, Dashboards, User and permission system
The crawler.
The technical website crawler was built entirely by Batunet. It captures complete websites including all subpages, runs technical checks and keeps the results historically, so changes remain visible over time. Crawls run automatically at regular intervals.
A run over a large website adds up to thousands upon thousands of URLs. That cannot happen in a single process. Processing is distributed across workers that pull their tasks from queues; on top of that sits custom job scheduling that prioritizes runs. Analysis, reporting and the interface on top were also built by Batunet.

When data volume becomes an architecture question.
Billions of records and a dataset in the terabyte range are not a metric but a constraint. At this scale, it is not the domain that decides where data lives, but which query must still be answerable in what time.
The consequence is the same everywhere: work that doesn’t have to happen in the request cycle doesn’t happen there. Distributed and asynchronous processing via workers and queues, caching, a search index, custom job scheduling and asynchronous generation of large exports are therefore part of the design, not add-ons.
Reporting and export.
Reports are individually configurable, generated automatically on a regular schedule and pull data together across multiple modules and data sources. Output is XLSX or CSV, including for large data exports.
None of this happens in the request cycle: generation runs asynchronously, and delivery follows automatically. The reporting system was built entirely by Batunet, the export system almost entirely.
External data sources.
The platform integrates numerous external APIs and data sources. That makes it dependent on systems that change without asking. Integration work here is not a by-product but an ongoing part of product development: connection, error handling, version changes, continuous technical maintenance.
We don’t name the providers and models behind it — neither for the data sources nor for the AI features.
AI in a mature product.
The AI feature set was built entirely by Batunet — not as a separate product alongside, but embedded in a system that already existed. It includes AI-assisted content features, visibility analysis in AI systems, analysis of AI answers and their sources, prompt processing, competitor analysis and automated recommendations.
The hard part wasn’t the model. It was the integration: existing data, existing permissions, existing workflows — and results that have to be reproducible enough to stand in a commercial product.
Technologies used
The stack was already in place when we took over, and in its fundamentals it matches what we build with anyway.
Where this case connects.
Questions about this case
Do you take over existing software that is already in production?
Yes. This case is exactly that: a running, commercially used product that we took over in 2017 and have developed ever since without a rebuild.
How do you handle large volumes of data?
This platform’s dataset is in the terabyte range, with billions of records. At this scale, storage is no longer a configuration question but part of the design: distributed and asynchronous processing, caching, a search index and custom job scheduling.
Do you only build the technology, or also the technical concept?
Both. The operator defines business requirements and product goals; the technical concept and implementation are largely developed by Batunet.
Why is no client name given here?
The case deliberately focuses on the engineering work. The client’s identity isn’t needed for that. We discuss named references in a direct conversation.
Continue your engineering journey.
Related concepts, decisions, playbooks and perspectives — as one connected path, not a list of links.
Technologies
Let’s talk about your project.
No sales team. A direct conversation with the management.
