Building a Multi-AI-Agent System for Software Testing to Save $4K per Engineer Monthly
A US-based company, Treegress, partnered with MobiDev to create a software testing automation solution based on multiple AI Agents.
The Story Behind the Multi-AI-Agent System for Software Testing
In the 2020s, the QA stage of software development takes up to 40% of the budget, and 80% of companies depend solely on manual testing. The need for testing automation is high. However, despite the marketing promise to significantly decrease testing time, the available testing automation tools have hidden time wastes.
The Treegress startup wanted to take a different approach to QA automation. In 2024, they partnered with MobiDev as they found our expertise in software development and testing invaluable. After the Tech consulting stage, it was suggested to build a system based on multiple AI Agents as it best fits the product goals.
Business Value of Building Multi-AI-Agent System for Software Testing
MobiDev built a multi-agent testing automation system that enables the QA teams not only to reduce the time and budget of software testing but also to make this process more efficient. This was achieved by the use of AI Agents that generate Test Cases and run them independently.
The key improvement metrics include:
- 8x faster regression cycles
- 30% fewer QA hours
- $2.5K-$4K saved per engineer per month
- 98% of QA runs successfully
Project Scope of Multi-AI-Agent System for Software Testing
MobiDev began cooperation with Treegress by providing Tech Consulting services in 2024. We worked closely with Treegress management to create a comprehensive Tech Strategy. We suggested an architecture that can independently analyze, plan, and execute testing, as well as assess the results in accordance with the tested product’s specifications.
In 6 months, we launched the MVP with visual testing capabilities. Following the initial product success, we continued our work on the core product features and added several more, including End-to-end functional testing from just a URL. As of now, we work closely with Treegress on supporting the stable operations of the product.
Deliverables for the Multi-AI-Agent System for Software Testing
MobiDev team built the multi-AI-agent QA system that has the following capabilities:
- Client-facing web application
- Website Analysis Engine
- AI Agent 1 that generates Test Cases and runs them independently
- AI Agent 2 that verifies the results of the test runs
- The Self-Healing Engine that uses the results of AI Agent 2 to improve the next testing
- Test Management System
- Chrome Extension
Tech Stack for Multi-AI-Agent System for Software Testing
Build Multi AI Agent System
Fill out the form and share your product vision. Our experts will get back to you within 1 business day.
FAQ
Yes, the architecture can be designed to connect with existing CI/CD pipelines, repositories, testing frameworks, issue trackers, and test management systems. The exact integration approach depends on the tools already used by the engineering team. Early integration planning helps ensure that the product supports existing workflows instead of creating an additional operational layer.
The system can ground test generation in product specifications, existing test documentation, application data, and predefined testing rules. A separate verification agent can review generated tests and evaluate execution results before they are accepted. Human approval can also be included for high-risk or business-critical scenarios.
Scalability should be considered at the architecture stage, particularly for test orchestration, parallel execution, job queues, data storage, and cloud infrastructure. Testing workloads can be distributed across multiple workers and scaled according to demand. Monitoring is also required to identify performance bottlenecks and control infrastructure and AI API costs.
Security measures should be selected according to the sensitivity of the data and the compliance requirements of the target customers. These may include encrypted data transfer and storage, access controls, isolated environments, secrets management, audit logging, and data-retention policies. The use of external AI providers should also be assessed to determine what information can be shared with third-party models.
Ownership terms should be defined clearly in the development agreement before the project begins. In a typical custom software engagement, the client receives ownership of the product-specific source code and deliverables created for the project. Any third-party technologies, open-source components, or external AI services remain subject to their respective licences and terms.
The system requires monitoring of test accuracy, agent behaviour, model performance, infrastructure costs, integrations, and execution reliability. Prompts, retrieval data, testing rules, and validation logic may also need to be updated as the tested products evolve. Regular maintenance helps prevent performance degradation and ensures that automated decisions remain aligned with current product requirements.
An existing product can first undergo a technical and product assessment to identify problems in its architecture, AI workflows, data, integrations, or business assumptions. Based on the findings, the team can recommend whether to improve individual components, rebuild the core system, or narrow the product scope. This assessment helps avoid investing further in an approach that cannot meet the required performance or business goals.
A multi-agent testing platform may require backend and frontend engineers, AI engineers, QA automation specialists, DevOps expertise, and product or project management. The exact team composition depends on the product scope and the maturity of the existing technology. The team can be adjusted as the product moves from technical discovery to MVP development and long-term scaling.
Enterprise readiness should be planned early rather than added after the product has already scaled. Important considerations include role-based access, tenant isolation, audit logs, deployment options, identity-provider integrations, data controls, service monitoring, and documented security processes. The specific requirements will depend on the industries and enterprise customers the startup intends to serve.
The architecture can potentially be extended to additional testing scenarios, but each type requires its own execution environment, tools, validation logic, and data. Mobile, API, performance, accessibility, and security testing may therefore be introduced as separate capabilities. It is usually more effective to validate one testing workflow first and expand the platform based on customer demand.