Testing Dynamics 365 used to be a project activity. You tested before go-live, before an upgrade, and before a big change, then the suite sat until the next one. That model no longer fits.
Three things have changed. Microsoft marked RSAT for deprecation, with support ending May 15, 2027. Microsoft now controls more of your testing calendar than most teams plan for. In Finance and Operations you can pause only one service update at a time, and a single update can change how the system behaves: in 10.0.49, more than 40 features became mandatory. In Business Central, an environment you do not schedule gets updated for you, and extensions that are not compatible may be uninstalled so the update can finish. You can defer an update for a while. You cannot skip validating it. And AI arrived on both sides of the test: AI tools now help write and maintain tests, and AI agents now act inside the ERP itself.
This guide is for ERP leaders, D365 architects, and QA and validation leads who need a testing program that holds up under all three. It covers what changed, why partners are moving testing into the implementation, the principles that still apply, where AI helps and where it should not decide, how to replace RSAT without losing what it did well, and a 90-day plan to get started.
The tooling, the roadmap, and the users of the system have all changed, on top of an update cadence that already made testing continuous. Each change has a direct consequence for how you test.
What Microsoft says: "marked for deprecation and won't be supported after May 15, 2027." The deprecation list names no replacement.
What it means for testing: Every RSAT user needs a new tool and a migration plan before May 2027.
What Microsoft says: Customers "can take up to four service updates per year and are required to take a minimum of two per year."
What it means for testing: Regression testing runs year-round.
What Microsoft says: Major updates in April and October, minor updates in other months, a five-month update period and a one-month grace period.
What it means for testing: You choose the update date. Validate early in the window instead of defaulting to the end.
What Microsoft says: Roadmap content publishes continuously starting September 2026. Release Planner retires November 15, 2026. Deployment schedules are unchanged.
What it means for testing: Review the roadmap on a set cadence, weekly or every two weeks.
What Microsoft says: Through MCP, agents can "perform nearly any function that's available to a user through the application interface."
What it means for testing: Agents are now users of your ERP. Their roles, permissions, and outcomes need testing too.
What Microsoft says: "Microsoft doesn't recommend one tool over another."
What it means for testing: The tool choice and the evaluation are yours.
The strongest implementation partners are moving testing to the start of the project. Partners we work with tell us testing is shifting from a phase before go-live into design, and that customers now expect it.
Microsoft's guidance points the same way. Its implementation guide says to start planning testing in the Initiate phase and to test "as early as possible," and warns against testing "too close to the go-live date."
The Business Process Catalog is what makes this practical:
When scope, requirements, and test cases share one catalog, three things change:
For partners, testing becomes a deliverable the customer keeps. For customers, the investment in testing during the project keeps paying back after it. Horizon is built around the Business Process Catalog for this reason.
These principles apply whatever tool you use, and AI makes each of them more important.
Microsoft's own implementation guide on testing strategy and its test strategy checklist are worth reading alongside these.
AI should do the work around the test. It should not be the judge of the test. Use models to create, maintain, scope, and triage. Keep execution and the pass or fail decision deterministic: the same test, on the same data and build, returns the same result every time.

Why draw the line at execution? Repeatability is what makes a result evidence. If a rerun can reach a different verdict, the result cannot serve as evidence. It also keeps cost predictable: a model in every execution path means a charge on every run. We cover this in more depth in Testing Needs Determinism.
Before you adopt any AI testing tool, get written answers to three questions:
For a broader governance frame, the NIST AI Risk Management Framework and, for life sciences, the FDA's Computer Software Assurance guidance (final, September 2025) both support a risk-based approach to assuring the tools you rely on.
A testing program that survives constant change has four parts. Most teams have the first and underinvest in the other three.
Coverage. Build a coverage map before you build tests. List your end-to-end processes (order to cash, procure to pay, record to report, plan to produce), rate each by business impact and change exposure, and mark where customizations and ISV solutions touch them. Standard processes still need coverage, but your customizations, integrations, and ISV seams are where only you can find the problem. Put most of your effort there.
Data. Decide where test data comes from, who owns it, and how it is refreshed after each environment copy. Keep it separate from test logic so business users can change a scenario without editing a test. Plan for the data a test consumes: a sales order test that ships the last unit of stock breaks the next run.
Be careful with AI-generated test data, especially in Finance and Operations. Meaningful data has to respect how your system is built: legal entities, number sequences, financial dimensions, posting profiles, and the relationships between customers, items, warehouses, and orders. A general-purpose model does not know that structure, so the records it creates are often in an invalid format or unrealistic for your business. AI is useful for masking real records into safe stand-ins and for simple generic values. For data that proves a process works, start from the structure of your own environment. That is how Horizon's TDM (Test Data Management) works: it builds data from the shape and structure of your Dynamics 365 environment, which Horizon discovers through Environment Scan, so the data fits your configuration.
Evidence. Every run should produce a record that links the result to the requirement or process it validates, the build and environment it ran on, and who approved it. Azure Test Plans links test cases to requirements and builds, and many teams already use it as their system of record. Whatever tool you use, keep that chain intact.
Cadence. For Finance and Operations, work backward from Microsoft's service update calendar:
For Business Central, validate in the preview window and pick your production date inside the five-month update period. Waiting until the end of the window leaves less time to fix what you find.
Most RSAT content explains why to leave. Before you replace it, write down what it did well, because your replacement should not be worse at any of it.
What RSAT got right. Per Microsoft Learn, RSAT lets "functional power users" record tasks in Task Recorder and turn them into automated tests "without needing to write source code." Test parameters "are decoupled from test steps and stored in Microsoft Excel files." And it is "fully integrated with Microsoft Azure DevOps for test execution, reporting, and investigation." Business ownership, separate data, and a traceable record: those are the three strengths to protect.
What it did not do. Recordings are brittle across updates, maintenance grows with every release, and coverage stops at Finance and Operations. Those are good reasons to move.
The options. Microsoft does not recommend one tool over another. Broadly, teams choose between:
We compare these paths in detail in Replacing RSAT: Choosing Between Playwright, AI, and Purpose-Built Test Automation.
Three questions for any replacement:
How to migrate. Start with priority, not volume. Sort your RSAT library into four groups: standard processes with no customization, customizations and ISV touchpoints, recordings that only check a screen loaded, and recordings you have re-recorded more than twice. The second group carries the most risk and comes first. The third group needs an outcome check added, and the fourth shows where maintenance has been costing you. Tools that import Task Recorder files and rebuild them automatically, as Horizon does, remove the rebuild effort, so this sorting sets priority and decides what to retire. Then run the old and new suites side by side through one service update and compare what each finds. Use that comparison as your acceptance test. Our RSAT migration guide walks through the steps.
Most AI testing content covers agents that test software. This section covers testing your ERP when agents work inside it. Through the Model Context Protocol for finance and operations apps, agents can "perform nearly any function that's available to a user through the application interface." Microsoft's overview of agents and Copilot in Dynamics 365 lists what is already available by app.
That changes the system under test in three ways.
Configuration now includes agent permissions. Microsoft is explicit that "the security role of the authenticated user for the agent determines which objects are returned." An agent with the wrong role can see or change what it should not. Test agent roles the way you test user roles: a matrix of what each should and should not be able to do, rerun after every security change and service update.
Verify outcomes, not paths. An agent may reach the right result by a different route each time. Instead of scripting its clicks, assert the business invariants that must hold however the work was done: journals balance, approval precedes payment, a purchase order never exceeds its budget, a customer on hold is never shipped to.
Guardrails are controls, so test them. A rule that says an agent cannot post above a threshold is a control like any other. If no test shows it holds after an update, you cannot show an auditor that it works.
For the agent's conversational behavior, Copilot Studio has its own tools: see Design a testing strategy for your agents and agent evaluation. Those tools grade answers by similarity and quality, which suits language. Your ERP still needs deterministic checks on what the agent actually did to the data. Horizon's agent validation is built for that second part.
You can move from a project-based suite to a standing program in one quarter. The goal of the quarter is to take one real service update through the new approach.

Track these from the first cycle:
If you want a second opinion on your coverage map or your RSAT library, talk to our team.
Microsoft: updates and deprecations
Microsoft: testing guidance and tools
Microsoft: implementation and the Business Process Catalog
Microsoft: agents and AI
Governance and regulation
From The TestMart

Real test scenarios. Real results. No sandbox demo.