October 1, 2026

The Dynamics 365 Testing Guide for the AI Era

A practical guide to Dynamics 365 testing after RSAT: what changed, where AI helps and where it should not decide, how to replace RSAT, and a 90-day plan.

The Dynamics 365 Testing Guide for the AI Era
Table of Contents
Book a Demo

Testing Dynamics 365 used to be a project activity. You tested before go-live, before an upgrade, and before a big change, then the suite sat until the next one. That model no longer fits.

Three things have changed. Microsoft marked RSAT for deprecation, with support ending May 15, 2027. Microsoft now controls more of your testing calendar than most teams plan for. In Finance and Operations you can pause only one service update at a time, and a single update can change how the system behaves: in 10.0.49, more than 40 features became mandatory. In Business Central, an environment you do not schedule gets updated for you, and extensions that are not compatible may be uninstalled so the update can finish. You can defer an update for a while. You cannot skip validating it. And AI arrived on both sides of the test: AI tools now help write and maintain tests, and AI agents now act inside the ERP itself.

This guide is for ERP leaders, D365 architects, and QA and validation leads who need a testing program that holds up under all three. It covers what changed, why partners are moving testing into the implementation, the principles that still apply, where AI helps and where it should not decide, how to replace RSAT without losing what it did well, and a 90-day plan to get started.

What changed

The tooling, the roadmap, and the users of the system have all changed, on top of an update cadence that already made testing continuous. Each change has a direct consequence for how you test.

RSAT deprecation

What Microsoft says: "marked for deprecation and won't be supported after May 15, 2027." The deprecation list names no replacement.

What it means for testing: Every RSAT user needs a new tool and a migration plan before May 2027.

F&O service update cadence

What Microsoft says: Customers "can take up to four service updates per year and are required to take a minimum of two per year."

What it means for testing: Regression testing runs year-round.

Business Central update timeline

What Microsoft says: Major updates in April and October, minor updates in other months, a five-month update period and a one-month grace period.

What it means for testing: You choose the update date. Validate early in the window instead of defaulting to the end.

Release waves replaced by the AI at Work roadmap

What Microsoft says: Roadmap content publishes continuously starting September 2026. Release Planner retires November 15, 2026. Deployment schedules are unchanged.

What it means for testing: Review the roadmap on a set cadence, weekly or every two weeks.

Agents in finance and operations apps

What Microsoft says: Through MCP, agents can "perform nearly any function that's available to a user through the application interface."

What it means for testing: Agents are now users of your ERP. Their roles, permissions, and outcomes need testing too.

Microsoft's tooling guidance

What Microsoft says: "Microsoft doesn't recommend one tool over another."

What it means for testing: The tool choice and the evaluation are yours.

Testing now starts in the implementation

The strongest implementation partners are moving testing to the start of the project. Partners we work with tell us testing is shifting from a phase before go-live into design, and that customers now expect it.

Microsoft's guidance points the same way. Its implementation guide says to start planning testing in the Initiate phase and to test "as early as possible," and warns against testing "too close to the go-live date."

The Business Process Catalog is what makes this practical:

  • It now carries test cases. The July 2026 release grew the catalog to 6,834 rows. Test cases grew from 1,082 to 1,628, including 316 new order-to-cash test cases.
  • It runs the project. The catalog imports into Azure DevOps as a project template, and the July release added a Business approval state field on requirements.
  • It starts at scoping. Since August 11, 2026, the Dynamics 365 Implementation Portal shows catalog process diagrams during project profiling.

When scope, requirements, and test cases share one catalog, three things change:

  1. Every process in scope has a test from day one. Coverage is decided at the same time as scope.
  2. Business sign-off attaches to a test. The process owner approves the test that proves the process works.
  3. The regression suite exists at go-live. Tests built during the implementation become the suite the customer runs on every service update.

For partners, testing becomes a deliverable the customer keeps. For customers, the investment in testing during the project keeps paying back after it. Horizon is built around the Business Process Catalog for this reason.

Seven principles that hold in the AI era

These principles apply whatever tool you use, and AI makes each of them more important.

  1. Test change, not screens. Scope every cycle from what actually changed: Microsoft's release notes, your code, ISV updates, configuration, feature flags, and data. Running the full suite on every update spends the window on processes that did not change.
  2. Let risk set coverage. Start with the processes that move money, move goods, or prove compliance. Microsoft's Business Process Catalog is a ready-made map for naming them and tracking which ones have tests.
  3. Keep test ownership with the people who know the process. The person who knows how a vendor invoice should post should define the test for it, whether or not they write code.
  4. Treat test data as part of the test. A test that fails because an item is on hold or a period is closed tells you nothing about the update. Valid data has to fit your configuration, so build it from your environment's structure, keep it separate from test logic, and version it.
  5. Produce evidence when the test runs. Auditors review your evidence of testing, not your testing. Screenshots and logs assembled after the fact are the weakest proof you can offer.
  6. Judge a suite by what it catches. A suite that has been green for a year may be testing processes that never change, or checking that screens load rather than that outcomes are correct. Track the last meaningful failure per test.
  7. Assume the environment moves under you. Feature management turns on new behavior, service updates autoupdate sandboxes before production, and ISVs ship on their own schedules. Your test plan has to know what state it is testing.

Microsoft's own implementation guide on testing strategy and its test strategy checklist are worth reading alongside these.

Where AI helps, and where it should not decide

AI should do the work around the test. It should not be the judge of the test. Use models to create, maintain, scope, and triage. Keep execution and the pass or fail decision deterministic: the same test, on the same data and build, returns the same result every time.

AI assists with creating, scoping, maintaining, and triaging tests. Execution, pass or fail, and evidence stay deterministic.
  • Turning a business process into test steps. What AI does: Drafts steps and checks from plain-language intent or a process description. Who decides: The process owner approves the test.
  • Scoping a service update. What AI does: Reads release notes and maps changes to the processes you run. Who decides: The QA or validation lead sets scope.
  • Creating test data. What AI does: Masks or varies real records, or fills simple generic values. Without your system's structure it cannot build meaningful Finance and Operations data on its own. Who decides: Data built from your environment's structure and configuration, versioned and reused.
  • Maintaining tests after UI changes. What AI does: Adapts tests automatically when Microsoft changes a form or control. Who decides: Every adjustment is logged for audit and review.
  • Triaging failures. What AI does: Sorts failures into data, environment, and product issues. Who decides: A person confirms every defect.
  • Executing the test and deciding pass or fail. What AI does: Not used. Who decides: Deterministic steps and assertions.
  • Producing evidence. What AI does: Summarizing results for readers. Who decides: The run itself generates the record.

Why draw the line at execution? Repeatability is what makes a result evidence. If a rerun can reach a different verdict, the result cannot serve as evidence. It also keeps cost predictable: a model in every execution path means a charge on every run. We cover this in more depth in Testing Needs Determinism.

Before you adopt any AI testing tool, get written answers to three questions:

  • Where does the model run, and what data leaves your tenant during creation, maintenance, and execution?
  • Can the same test produce a different verdict on a rerun?
  • What does a run cost when a model is involved, and who pays for retries?

For a broader governance frame, the NIST AI Risk Management Framework and, for life sciences, the FDA's Computer Software Assurance guidance (final, September 2025) both support a risk-based approach to assuring the tools you rely on.

Building the program: coverage, data, evidence, cadence

A testing program that survives constant change has four parts. Most teams have the first and underinvest in the other three.

Coverage. Build a coverage map before you build tests. List your end-to-end processes (order to cash, procure to pay, record to report, plan to produce), rate each by business impact and change exposure, and mark where customizations and ISV solutions touch them. Standard processes still need coverage, but your customizations, integrations, and ISV seams are where only you can find the problem. Put most of your effort there.

Data. Decide where test data comes from, who owns it, and how it is refreshed after each environment copy. Keep it separate from test logic so business users can change a scenario without editing a test. Plan for the data a test consumes: a sales order test that ships the last unit of stock breaks the next run.

Be careful with AI-generated test data, especially in Finance and Operations. Meaningful data has to respect how your system is built: legal entities, number sequences, financial dimensions, posting profiles, and the relationships between customers, items, warehouses, and orders. A general-purpose model does not know that structure, so the records it creates are often in an invalid format or unrealistic for your business. AI is useful for masking real records into safe stand-ins and for simple generic values. For data that proves a process works, start from the structure of your own environment. That is how Horizon's TDM (Test Data Management) works: it builds data from the shape and structure of your Dynamics 365 environment, which Horizon discovers through Environment Scan, so the data fits your configuration.

Evidence. Every run should produce a record that links the result to the requirement or process it validates, the build and environment it ran on, and who approved it. Azure Test Plans links test cases to requirements and builds, and many teams already use it as their system of record. Whatever tool you use, keep that chain intact.

Cadence. For Finance and Operations, work backward from Microsoft's service update calendar:

  1. Preview. Update a sandbox and run a smoke test on money and goods processes. Ask ISVs for compatibility confirmation in writing.
  2. General availability. Run scoped regression: full coverage on critical paths, targeted coverage where the release notes touch your modules. Review features turned on by default.
  3. Sandbox autoupdate. Sandboxes update seven days before production. Treat that week as your final check.
  4. Production. Choose the date by your business calendar, not the default. If you need more time, you can pause one update at a time, but you must still take at least two per year.

For Business Central, validate in the preview window and pick your production date inside the five-month update period. Waiting until the end of the window leaves less time to fix what you find.

Replacing RSAT: what to keep, what to evaluate

Most RSAT content explains why to leave. Before you replace it, write down what it did well, because your replacement should not be worse at any of it.

What RSAT got right. Per Microsoft Learn, RSAT lets "functional power users" record tasks in Task Recorder and turn them into automated tests "without needing to write source code." Test parameters "are decoupled from test steps and stored in Microsoft Excel files." And it is "fully integrated with Microsoft Azure DevOps for test execution, reporting, and investigation." Business ownership, separate data, and a traceable record: those are the three strengths to protect.

What it did not do. Recordings are brittle across updates, maintenance grows with every release, and coverage stops at Finance and Operations. Those are good reasons to move.

The options. Microsoft does not recommend one tool over another. Broadly, teams choose between:

  • Open-source frameworks such as Playwright. Microsoft's FastTrack team has published a Finance and Operations Playwright example. No license fee, but you fund the engineering and the maintenance. Microsoft also deprecated Power Apps Test Engine in April 2026 and now points Power Platform teams to Playwright directly.
  • Purpose-built platforms designed for Dynamics 365, which trade a subscription for a lower technical burden.
  • AI-assisted tools, which speed up creation and maintenance. Apply the questions in the section above before you trust one with pass or fail.

We compare these paths in detail in Replacing RSAT: Choosing Between Playwright, AI, and Purpose-Built Test Automation.

Three questions for any replacement:

  1. Can a functional user create and change a test without an engineer?
  2. Is test data managed separately from test logic, and can the business edit it?
  3. Does every run produce evidence linked to the requirement it validates, in a system your auditors already accept?

How to migrate. Start with priority, not volume. Sort your RSAT library into four groups: standard processes with no customization, customizations and ISV touchpoints, recordings that only check a screen loaded, and recordings you have re-recorded more than twice. The second group carries the most risk and comes first. The third group needs an outcome check added, and the fourth shows where maintenance has been costing you. Tools that import Task Recorder files and rebuild them automatically, as Horizon does, remove the rebuild effort, so this sorting sets priority and decides what to retire. Then run the old and new suites side by side through one service update and compare what each finds. Use that comparison as your acceptance test. Our RSAT migration guide walks through the steps.

Testing agents and Copilot inside Dynamics 365

Most AI testing content covers agents that test software. This section covers testing your ERP when agents work inside it. Through the Model Context Protocol for finance and operations apps, agents can "perform nearly any function that's available to a user through the application interface." Microsoft's overview of agents and Copilot in Dynamics 365 lists what is already available by app.

That changes the system under test in three ways.

Configuration now includes agent permissions. Microsoft is explicit that "the security role of the authenticated user for the agent determines which objects are returned." An agent with the wrong role can see or change what it should not. Test agent roles the way you test user roles: a matrix of what each should and should not be able to do, rerun after every security change and service update.

Verify outcomes, not paths. An agent may reach the right result by a different route each time. Instead of scripting its clicks, assert the business invariants that must hold however the work was done: journals balance, approval precedes payment, a purchase order never exceeds its budget, a customer on hold is never shipped to.

Guardrails are controls, so test them. A rule that says an agent cannot post above a threshold is a control like any other. If no test shows it holds after an update, you cannot show an auditor that it works.

For the agent's conversational behavior, Copilot Studio has its own tools: see Design a testing strategy for your agents and agent evaluation. Those tools grade answers by similarity and quality, which suits language. Your ERP still needs deterministic checks on what the agent actually did to the data. Horizon's agent validation is built for that second part.

A 90-day plan and the metrics that matter

You can move from a project-based suite to a standing program in one quarter. The goal of the quarter is to take one real service update through the new approach.

  1. Days 1 to 30: map. Build the coverage map of your critical processes and mark customizations, integrations, and ISV touchpoints. Name an owner for each process. If you are mid-implementation, do this with your implementation partner during design, starting from the Business Process Catalog. If you use RSAT, triage the library into the four groups above. Decide where evidence will live.
  2. Days 31 to 60: choose and prove. Shortlist tools against the three RSAT questions and the three AI questions. Pilot on three to five critical processes, including at least one customization and one integration. Set up test data that business users can edit.
  3. Days 61 to 90: run a real cycle. Take a service update through the new program from preview to production. For Finance and Operations, 10.0.50 enters preview October 23, 2026 and is generally available December 22, 2026, which lines up with a fourth-quarter start. Run the old and new suites side by side and compare findings.
A 90-day plan for Dynamics 365 testing: map in days 1 to 30, choose and prove in days 31 to 60, run a real cycle in days 61 to 90.

Track these from the first cycle:

  • Share of critical processes with automated coverage: whether coverage follows risk.
  • Share of tests on customizations and ISV touchpoints: whether effort is where only you can find problems.
  • Days from update availability to validated: how much of the update window you actually use.
  • Defects found before production, per cycle: whether the suite catches anything.
  • Maintenance hours per update: whether the suite is getting cheaper or more expensive to keep.
  • Time since last meaningful failure, per test: which tests have stopped earning their place.

If you want a second opinion on your coverage map or your RSAT library, talk to our team.

Resources

Microsoft: updates and deprecations

Microsoft: testing guidance and tools

Microsoft: implementation and the Business Process Catalog

Microsoft: agents and AI

Governance and regulation

From The TestMart

Daniel Diefendorf
CEO

See it in your environment, today.

Real test scenarios. Real results. No sandbox demo.