Back to Blog AI
9.25.2026

What happens when your ERP's AI agent approves a bad invoice

By Meredith Grace

Senior Manager Content Marketing at Medius


An invoice gets coded to the wrong account, clears matching, and gets paid. That has happened in every AP department that has ever existed. What is new is the answer to the next question. Who made the call?

If the answer is an AI agent that shipped with your ERP, the follow-up questions get harder quickly. On what basis did it decide? Who owns the outcome? And can you reconstruct the reasoning for an auditor 14 months later, when the person who configured the agent has changed jobs?

Those questions are no longer hypothetical. Oracle, Microsoft, and SAP have all put AI into the payables process, and they have not drawn the autonomy line in the same place. Most finance organizations have not written down where they think the line belongs at all.

The distinction matters: ERP-native agents create accountability gaps that governed AP automation is specifically designed to close. Understanding where those gaps open is the first step to knowing why the design of your AP layer matters as much as the AI it uses.

What the ERP vendors actually shipped, and where each drew the line

The differences between them matter more than the category label does.

Oracle's Payables Agent, part of Fusion Cloud ERP, is built for what Oracle calls near-touchless processing. It uses large language models to interpret invoice documents, extract and map data to invoice attributes, autofill fields, and carry invoices through the workflow toward payment readiness. People enter on exceptions. An invoice gets flagged, and a person cancels it, corrects the captured data, or releases the exception if they think the flag was wrong.

Microsoft drew a harder line. The Business Central Payables Agent monitors a payables inbox, extracts invoice data, matches vendors through both exact and fuzzy matching, and suggests general ledger classifications. Then it stops. Microsoft's own documentation states that the agent "never automatically posts invoices or makes permanent changes to your financial data without explicit human approval." Creating a vendor and converting a draft into a purchase invoice are both classified as high-risk actions requiring mandatory human sign-off.

SAP has been more conservative again, at least in payables. Joule operates as an assistant for AP work rather than an autonomous approver, and SAP's named finance agents released through the first quarter of 2026 sit in adjacent territory: a Dispute Resolution Agent in beta, plus expense validation and pre-submit audit agents. No shipped SAP agent approves supplier invoices on its own.

So agentic AI in payables is not one thing. One vendor is optimizing for straight-through processing, another has engineered mandatory stops into the workflow, and a third is still mostly assisting. A control policy written for "AI agents" as a single category will fit at least one of those badly.

The accountability gap opens in the middle, not at full autonomy

Full autonomy is the easy case to govern. If an agent processes an invoice end to end with no human involved, ownership is unambiguous, and most organizations would build a control around it precisely because the absence of a person is so obvious.

The hard case is the hybrid, which is also the common one. According to Ardent Partners research, the average enterprise processes 32.6% of invoices straight through, and even best-in-class organizations only reach 49.2%. Most invoices still pass in front of a person.

That person is increasingly approving a decision an agent already made. The agent read the document, picked the vendor, chose the account, and presented a finished draft. The approval click is real. The review often is not. When something goes wrong, the organization has a record showing that a named employee approved the invoice, and that employee will say the system recommended it. Neither statement is false, and neither one settles anything.

A log is not the same thing as accountability. Most of these platforms produce a detailed record of what happened, and that record will tell you the agent extracted a vendor name at 11:42 and a person approved the draft at 14:15. What it will not tell you is who was responsible for the judgment in between. That question does not have a technical answer. It has an organizational one, and it needs to be settled before the invoice goes wrong rather than during the post-mortem.

The failure mode is also worse than it first sounds. An approver working through 60 invoices is not looking for a blank field, which is easy to spot. They are looking at a complete, internally consistent, entirely plausible draft that happens to name the wrong entity in a group of similarly named suppliers. Extraction models fail by producing confident answers, not obvious gaps, and confident wrong answers are precisely the ones an agent will not escalate for review.

The data behind AP's accountability problem

Ardent Partners surveyed hundreds of finance leaders on how AI is reshaping the payables process, and where control is getting harder to maintain. Read the full report to see where your peers are drawing the line.

Read the report

Segregation of duties assumes the approver is a person

Segregation of duties exists to keep whoever creates a transaction from being whoever approves it. The control assumes two parties with independent judgment and different incentives.

An agent that codes an invoice and routes it to a human approver has not broken that rule on paper. But if the same agent also determines who receives the approval, applies the threshold logic, and pre-fills the decision, the independence is thinner than the control diagram suggests. When one agent identity performs steps that used to be split between an AP clerk and an AP manager, duties have been consolidated without anyone documenting the change. Approval workflows already create fraud risk when they are poorly designed. Adding an actor that touches several steps at once does not simplify that picture.

AI agent

Microsoft's design points at the right instinct here. The Business Central agent operates under its own unique user identity, every action it takes is attributed to that identity, and administrators assign the specific permission sets that define what it can reach. Treat the agent as a user, give it a role, and constrain it like you would constrain anyone else. Organizations that let an agent act through a shared service account, or under the credentials of whoever set it up, give up the ability to answer basic questions later.

What audit-ready evidence for an AI decision looks like

An auditor asking about an AI-approved invoice is not asking whether you used AI. They are asking you to show the decision. In practice, that means six things:

Attribution

Which identity took each action, agent or human, at every step

Inputs

What the agent actually saw, including the document, the vendor record, the purchase order, and prior invoice history

Reasoning

Why it chose this vendor match and this account, captured when the decision was made rather than reconstructed afterward

Uncertainty

Whether the agent was confident, and what it did when it was not

Human action

What the reviewer was shown and what they changed, recorded separately from what the agent proposed

Retention

All of the above, available for as long as your audit and record retention obligations run

That last one catches people out. Microsoft retains agent prompts and responses for 20 days in Business Central for support purposes. That is a diagnostic tool, not an audit archive, and it is a reasonable design choice on Microsoft's part. It only becomes a problem when the evidence needed to reconstruct a decision exists nowhere else. Twenty days is a fraction of the window in which a dispute, a fraud investigation, or an external audit will surface the question.

Governance has to sit outside the agent

There is a structural problem with asking the agent's own platform to be the record of what the agent did. If the system making the decision is also the system storing the evidence and enforcing the control, then the quality of your control is whatever the vendor decided to build, and it changes when they change it. Independence is the whole point of a control, and a closed loop does not have any.

This is the case for keeping invoice governance in a layer that sits outside the ERP and outside any single agent. Capture, coding, three-way matching, approval routing, and fraud screening all run against the same rules in Medius AP Automation, and they produce the same audit trail, whether an invoice arrived by email, through e-Invoicing, or with coding already proposed by an ERP-native agent upstream. The evidence lives in one place and follows one retention policy, rather than being scattered across whichever tools happened to touch the invoice.

It also means AI can be applied where it helps without becoming the system of record for its own behavior. Medius Copilot gives approvers context on an invoice before they approve it, which is the opposite of a one-click rubber stamp, and Medius AI operates inside the same governed process rather than beside it.

The agents are going to keep getting more capable, and the pressure to let them run further will keep increasing. That is a reasonable direction. But the question an auditor asks after a bad invoice has not changed at all: who approved this, on what basis, and where is that written down? It is much cheaper to be able to answer that before the question gets asked.

Talk to us about your AP governance layer

Medius keeps capture, matching, approval routing, and fraud screening in a single auditable layer, outside any ERP agent, and outside any single vendor's retention policy. See how it works for your process.

Get in touch

two women looking at computer screen line art

Frequently asked questions

Technically yes, depending on the platform and how it is configured. Oracle's Payables Agent is built for near-touchless straight-through processing, while Microsoft's Business Central agent is explicitly designed not to post invoices without human approval. The more useful question is whether your own configuration allows it, which is a policy decision rather than a product limitation.

It depends entirely on the vendor. Microsoft states its Payables Agent "never automatically posts invoices or makes permanent changes to your financial data without explicit human approval," and treats vendor creation as a high-risk action requiring sign-off. Oracle's agent is designed to move invoices through to payment readiness with people handling flagged exceptions. Check the specific product documentation rather than assuming.

The organization is. No current ERP agreement transfers financial liability for a bad payment to the software vendor, and an agent is not a legal actor. Internally, accountability usually lands on whoever owns the control, which is why it matters that someone is formally named. Absent that, a bad invoice tends to produce a dispute between AP, IT, and the person who clicked approve.

Enough to reconstruct the decision. That means the identity that acted, the inputs the agent used, the reasoning behind the coding or vendor match, any uncertainty it flagged, and what the human reviewer actually did. Screenshots and summary logs generally will not satisfy this. The evidence has to be captured when the decision happens, not assembled later.

Not automatically, but it can erode it quietly. Problems arise when one agent identity performs steps that were previously split between roles, when the agent both proposes a decision and selects its approver, or when it operates through a shared account rather than its own. Giving the agent a distinct identity with defined permissions preserves the control.

Yes. A dedicated identity means every action is attributed to the agent rather than a person or a generic service account, permissions can be scoped precisely, and access can be revoked without disrupting anyone's job. Microsoft built this into the Business Central Payables Agent. Where a platform does not provide it, it is worth engineering around.

Treat the agent as an in-scope actor in the control narrative rather than as a tool. Document what it is permitted to do, what requires human approval, how exceptions escalate, and where the decision evidence is retained and for how long. Control testing then covers the agent's actions the same way it covers a person's.

Medius Financial Census 2026

87% of finance pros have ignored suspected fraud. That's just one finding. See what 2,386 finance leaders revealed about fraud, late payments, AI, and the profession's future.

Get the report

Ardent Partners' State of AP in 2026

AI is changing what AP teams can do, and fast. Get leading analysts at Ardent Partners' take on where the industry is headed.

Get the report

Webinar: Preparing for agentic AI in AP

In this webinar with SSON, see how leading finance teams are moving from automation to autonomous AP, with real governance built in.

Watch the webinar

Discover accounts payable benchmarks

Learn the efficiency metrics that matter for AP teams and the benchmarks derived from thousands of Medius customers around the globe.

Get the report

Questions about AP automation?

Let's talk through it. A 30-minute conversation with a Medius expert costs you nothing and might save you a lot.

Book a consult

Ready to transform your AP? 

Book a Demo Contact Us