The convergence of generative artificial intelligence and automated software testing has ushered in a new era of "agentic engineering," a shift that demands a fundamental recalibration of how developers approach code quality and system reliability. On September 25, 2026, David Burns, Head of Developer Advocacy and Open Source at BrowserStack, joined the Stack Overflow Podcast to discuss the critical role of professional skepticism in an AI-saturated landscape. As software development cycles accelerate under the influence of Large Language Models (LLMs), the industry faces a dual challenge: leveraging the productivity gains of autonomous agents while mitigating the inherent non-determinism of AI-generated outputs. Burns, a veteran in the testing space, emphasized that the solution lies not in abandoning traditional rigor, but in adapting classic methodologies—such as test-driven development (TDD)—to the unique demands of agentic workflows.
The Rise of Agentic Engineering and the Necessity of Skepticism
The term "agentic engineering" refers to the transition from using AI as a simple autocomplete tool to deploying autonomous agents capable of planning, executing, and refining complex software tasks. While these agents can significantly reduce the cognitive load on human developers, they also introduce a layer of opacity into the development process. Burns argued that professional skepticism is no longer just a soft skill but a technical requirement for modern engineers. This skepticism involves questioning the "hallucinations" of AI agents and ensuring that the logic produced aligns with the intended business requirements.
Industry data supports this cautious approach. According to a 2026 State of DevOps report, while 78% of enterprise organizations have integrated AI agents into their CI/CD pipelines, nearly 45% have reported an increase in "silent failures"—bugs that do not trigger traditional errors but result in incorrect data processing or logic errors. Burns noted that the reliance on AI-generated code without rigorous verification can lead to a degradation of the codebase over time, a phenomenon often referred to as "technical debt inflation."
To combat this, Burns suggests that developers must maintain a "testing-first" mindset. This involves defining the expected outcome of a task before the AI agent begins its work. By setting clear boundaries and assertions, developers can create a safety net that catches anomalies before they reach production environments.
Applying Test-Driven Development to AI Agents
One of the more provocative insights shared by Burns is the application of Test-Driven Development (TDD) to the engineering of agents. Traditionally, TDD follows a "red-green-refactor" cycle: write a failing test, write the minimum code to pass the test, and then clean up the code. When applied to agentic engineering, this framework ensures that the agent’s behavior is constrained by pre-defined requirements.
In an agentic context, TDD serves as a specification for the agent’s objective. Instead of merely asking an AI to "build a login feature," a developer using TDD would first write a series of automated tests covering edge cases, such as SQL injection attempts, expired tokens, and multi-factor authentication failures. The agent is then tasked with generating code that satisfies these specific tests. This approach shifts the developer’s role from a primary coder to a "quality architect" who defines the success criteria that the AI must meet.
This methodology addresses one of the primary criticisms of AI in software development: the lack of reproducibility. By anchoring agentic behavior in a suite of deterministic tests, organizations can reclaim control over their software’s stability. Burns highlighted that BrowserStack has been at the forefront of this movement, providing the infrastructure necessary to run these complex, multi-layered tests at scale.
The Persistent Challenge of Flaky Tests and Application State
Beyond the challenges of AI, the conversation addressed a perennial issue in software engineering: flaky tests. A flaky test is defined as a test that provides both passing and failing results without any changes to the underlying code. These tests are the "silent killers" of developer productivity, as they erode trust in automated testing suites and cause unnecessary delays in the deployment pipeline.
Burns posited that the root cause of flakiness is almost always related to the management of application state. Software applications are complex state machines; if a test does not account for the exact state of the database, the network, or the browser environment, it is prone to failure. "Fixing flaky tests isn’t about re-running them until they pass," Burns noted during the discussion. "It’s about isolating the state and ensuring that every test starts from a clean, known baseline."
Data from BrowserStack’s internal telemetry indicates that nearly 60% of flaky tests in modern web applications are caused by asynchronous operations and race conditions. When a test attempts to interact with a UI element before the backend state has been fully updated, the result is an inconsistent failure. Burns advocated for better "state synchronization" techniques, where the testing framework is deeply integrated with the application’s lifecycle, rather than merely acting as an external observer.
Chronology of the Testing Evolution
The insights shared by Burns reflect a broader historical shift in the software industry over the last decade. To understand the current state of agentic engineering, it is necessary to look at the timeline of testing evolution:
- 2010-2015: The Automation Era. The industry moved away from manual "point-and-click" testing toward automated scripts using tools like Selenium. David Burns, as a core contributor to the Selenium project, witnessed the initial struggle to make automated tests reliable.
- 2016-2021: The CI/CD Integration Phase. Testing became a mandatory part of the "shift-left" movement. Tools like BrowserStack allowed developers to test across thousands of device and browser combinations in the cloud, making cross-platform compatibility a standard requirement.
- 2022-2024: The Generative AI Explosion. The release of advanced LLMs led to the widespread adoption of AI coding assistants. Testing began to focus on validating AI-generated snippets, but the process remained largely human-led.
- 2025-2026: The Age of Agents. The industry entered the current phase, where AI agents act autonomously. This period is defined by the need for "agentic orchestration" and the rigorous application of TDD to non-human developers.
Supporting Data on Developer Productivity and Quality
The implications of these shifts are measurable. A 2026 survey of 2,500 software engineers conducted by a leading technology research firm found that developers spend an average of 35% of their time managing "test debt," which includes fixing flaky tests and updating outdated test suites. Organizations that have successfully implemented agentic TDD reported a 22% reduction in time-to-market for new features, alongside a 15% increase in production stability.
Furthermore, the role of open source in this ecosystem cannot be overstated. Burns, who leads the open-source office at BrowserStack, emphasized that the tools used to test AI must themselves be transparent and community-vetted. The reliance on proprietary AI "black boxes" for testing purposes creates a single point of failure that many in the engineering community find unacceptable.
Official Responses and Community Impact
The developer community has reacted with a mixture of optimism and caution to the concepts discussed by Burns. On platforms like Stack Overflow, discussions regarding "AI-driven TDD" have seen a 140% increase in engagement over the past twelve months. Many senior architects agree that the "skeptical engineer" is the most valuable asset in a modern dev shop.
Stack Overflow continues to reward this culture of rigorous knowledge sharing. For instance, the community recently recognized user "brentvatne" with a Populist badge for providing a highly effective answer on how to read app.json or exp.json files programmatically. This type of foundational knowledge—understanding how to manipulate configuration and state—remains essential even as higher-level tasks are handed off to AI agents.
Broader Impact and Future Implications
The shift toward agentic engineering and the renewed focus on state management will likely redefine the software engineering curriculum. Future developers will need to be less focused on syntax and more focused on "systemic verification." The ability to architect a testable system will become the primary differentiator between junior and senior talent.
For companies like BrowserStack, the mission is clear: provide the "ground truth" for software behavior. As AI agents become more prevalent, the demand for a platform that can provide a deterministic, real-world environment for verification will only grow. The future of software quality lies in the balance between the creative speed of AI and the disciplined skepticism of the human engineer.
In conclusion, the insights provided by David Burns serve as a roadmap for navigating the complexities of 2026’s technological landscape. By treating AI agents as powerful but fallible tools, and by applying rigorous testing frameworks like TDD to their outputs, the industry can harness the power of artificial intelligence without sacrificing the reliability that users demand. The management of application state remains the final frontier in the battle against flakiness, requiring a deep understanding of the underlying mechanics of modern web and mobile applications. As the industry moves forward, the "skeptical professional" will remain the ultimate safeguard against the unpredictability of an AI-driven world.








