[seopress_breadcrumbs]

Building Governed Agentic AI for Financial Operations

Emerj | Reindeer | JPMorganChase | AI in Financial Services

 

This article is sponsored by Reindeer and was written, edited, and published in alignment with our Emerj sponsored content guidelines. Learn more about our thought leadership and content creation services on our Emerj Media Services page.

Financial institutions face a severe operational bottleneck: core workflows rely on fragmented, unstructured data, require nuanced human judgment, and operate under strict regulatory scrutiny. As detailed in a report published by the Bank for International Settlements, expanding AI integration across financial services introduces major operational, model, and data governance risks when systems are deployed without standardized controls.

BFSI environments ingest disparate document formats across core databases, risk engines, and document repositories. The CFA Institute reported that 90% of enterprise data is unstructured. Data heterogeneity without governance increases processing error rates in document-intensive workflows like Anti-Money Laundering (AML) and Know Your Customer (KYC).

Autonomous decision engines that lack explainability conflict with financial compliance mandates requiring traceable decision lineage. Research highlighted by the ProSight Financial Association demonstrated that non-compliance costs institutions an average of $14.82 million annually, which is 2.71 times the cost of maintaining compliance infrastructure.

Financial automation initiatives encounter friction when transitioning from testing environments to live production. Analysis published by the IEEE Computer Society indicated that while 83% of technology leaders initiate AI projects, only 9% successfully operationalize them, resulting in stalled deployments as policies and underlying systems evolve.

Full autonomy remains limited when handling non-standard workflows. Research published by MR Online found that current AI task execution achieves an average success rate of 30% for complex end-to-end workplace processes. Guidelines published by the National Institute of Standards and Technology emphasize that reliable deployment requires continuous human oversight and explicit fallback mechanisms when model confidence declines.

Emerj’s Yolandi de Weerdt hosted conversations with the Co-Founder and Co-CEO of Reindeer, Yoav Naveh, and Ajay Swamy, Senior Executive Product Director – GenAI Products, AIML Platform Management and Governance at JPMorganChase, to clarify how leading institutions are transitioning from isolated AI experiments to operationally governed agentic systems and the structural requirements for deploying them safely inside complex, regulated financial workflows.

This article examines four operational insights that matter most for BFSI leaders working to deploy agentic AI into core financial workflows safely:

  • Workflow redesign for automation‑ready operations: Restructure end‑to‑end processes so agents can execute across fragmented systems, inconsistent documents, and human‑judgment checkpoints without breaking under the weight of exceptions.
  • Governance‑first control for explainable agent decisions: Enforce auditability at every decision point so agent outputs carry traceable lineage, policy context, and defensible reasoning that compliance teams can surface instantly.
  • Exception‑aware escalation for regulated workflows: Equip agents with structured unknown detection so ambiguity, edge cases, and policy conflicts trigger controlled human intervention instead of silent failure or hallucinated output.
  • Strategic operating model for sustainable agent deployment: Assign long‑term ownership for agent maintenance, versioning, and governance so prototypes evolve into durable operational systems rather than accumulating tech debt.

Listen to the full episodes below:

Episode 1:  Managing AI Agents at Scale Across BFSI Operations – with Yoav Naveh of Reindeer AI

Expertise: AI Transformation, AI Agents, Enterprise Automation, Process Mining

Brief Recognition: Yoav Naveh is Founder and CEO of Reindeer, following a career spanning technology entrepreneurship, operations, and venture investing. He co-founded ConvertMedia and scaled the company to a $50M run rate before its acquisition by Taboola for nearly $100M, where he later held leadership roles in video and people operations. He is also Managing Partner at INT3, an investment firm focused on early-stage startups, and previously served in Unit 8200 of the Israel Defense Forces, contributing to the development of its early cyber capabilities. He holds a B.Sc. in Mathematics and Computer Science from Tel Aviv University.

Episode 3:  Managing Change Across Modern Financial Operations – with Ajay Swamy of JPMorganChase and Founder of FundLens.ai

Expertise: GenAI Product Strategy, AI/ML Platforms, AI Governance, Product & Technology Leadership

Brief Recognition: Ajay Swamy is Senior Executive Director of GenAI Products, AI/ML Platform Management and Governance at JPMorganChase and Chief AI Officer at FundLens.ai. Previously, he spent nearly five years at AWS, most recently leading worldwide cross-industry GenAI solutions, and held product leadership roles at McKinsey, where he developed ML-powered SaaS products generating $10M+ in revenue, including solutions for financial services and manufacturing. He also co-founded and exited Client Pay Direct, a B2B2C fintech platform, after building it to $500K in ARR in its first year. He holds an MBA from IE Business School.

Workflow Redesign for Automation‑Ready Operations

“It’s no longer just a question and answer. You actually have to take a task across multiple systems, multiple teams, and it’s a lot more difficult to just close a task. Exceptions and edge cases become much more frequent.”  

Yoav Naveh, Co‑Founder and Co‑CEO at Reindeer

Yoav Naveh starts the series by reframing workflow automation as a structural challenge rather than a capability challenge. He points out that BFSI institutions don’t struggle because any individual step is complex — they struggle because the moment ten simple steps span five systems, three teams, and inconsistent document formats, the work becomes something entirely different.

​According to Naveh, leaders must treat workflow redesign as an operational engineering problem: map the real path work takes, surface where judgment actually lives, and stabilize exception handling before introducing any autonomous execution.

Ajay Swamy adds the governance dimension to this redesign in his conversation with Emerj. He argues that workflows fail not at scale, but at the seams — where fragmented data, inconsistent entitlements, and shifting policy interpretation collide. For Ajay, automation‑ready workflows require institutions to expose every system the process touches, every data shape it consumes, and every point where human judgment modifies the outcome. Without that visibility, any attempt at automation accelerates the underlying fragmentation.

Their guidance forms a practical sequence leaders can use to redesign workflows, so agentic AI can operate without collapsing under exceptions:

  • Start with the workflow exactly as humans run it today: Naveh stresses that the first version must mirror the current process, not an imagined future one.
  • Digitize the workflow end‑to‑end before attempting any optimization: This creates a stable baseline and exposes hidden dependencies across systems and teams.
  • Map every system, data store, and judgment checkpoint the workflow touches: Ajay highlights that fragmentation — not volume — is the real operational barrier.
  • Stabilize exception handling as a first‑order design requirement: Naveh notes that exceptions multiply in multi‑system execution, and agents must be designed to survive them.
  • Only after the baseline is stable, introduce new capabilities or reimagine the workflow: Once digitized, institutions can safely layer in enhancements such as public‑record enrichment or dynamic document validation.

This sequence reflects the core message from both guests: automation‑ready workflows are engineered, not discovered. Institutions must build the operational foundation first — the systems map, the judgment map, the exception architecture — before agentic AI can execute reliably inside regulated financial operations.

Governance‑First Control for Explainable Agent Decisions

Ajay Swamy pushes the conversation into governance, arguing that it is not a safeguard around agentic AI, but rather the condition that determines whether the technology can be used at all. He argues that institutions routinely misjudge the risk surface by focusing on model capabilities rather than supervisory controls.

Ajay Swamy frames the governance challenge in the clearest possible terms:

“A number or a decision that you cannot trace is a number or a decision that you cannot defend. You’re not really evaluating the technology; the technology is the easy part because the demos all look great. What you’re evaluating is whether you can actually supervise it and explain every step end‑to‑end.”

– Ajay Swamy, Senior Executive Product Director – GenAI Products, AIML Platform Management and Governance at JPMorgan Chase

Swamy’s point is that explainability is not a compliance preference; it is the operational backbone of regulated AI. Every autonomous action must carry lineage, policy context, and a defensible chain of reasoning that can be surfaced instantly — not reconstructed after the fact. He stresses that institutions should treat explainability as a design constraint, not a reporting requirement, because the moment an agent acts without traceability, the institution inherits regulatory exposure it cannot mitigate.

Yoav Naveh adds the operational counterpart to Swamy’s governance stance by noting that institutions often assume accuracy is the primary risk, when the real exposure emerges when an agent encounters uncertainty or as work inevitably changes. According to Naveh, governance must be engineered to detect unknowns early and route them into structured human intervention. His emphasis is that agents must be built to escalate ambiguity rather than mask it, because silent failure is far more dangerous than an explicit handoff.

The combined takeaway is that governance‑first design is not a philosophy; it is a sequence of supervisory conditions that must be met before any agent is allowed to operate within a regulated workflow. Institutions must be able to:

  • Surface the full lineage of every decision, including data inputs, transformations, and policy interpretation.
  • Identify where human judgment re-enters the workflow, and ensure those checkpoints are explicit rather than implied.
  • Detect divergence in real time — not through post‑hoc reconstruction after an incident.
  • Govern agent evolution through controlled versioning, regression testing, and approval cycles.

Swamy’s contribution defines the standard: if a decision cannot be explained, it cannot be defended. Naveh’s contribution defines the mechanism: agents must be designed to recognize uncertainty and escalate it predictably. They establish the core principle of this subsection — agentic AI only becomes viable in BFSI when governance is engineered as the first layer, not the last.

Exception‑Aware Escalation for Regulated Workflows

Yoav shifts the conversation from governance to the operational reality of what happens when an agent reaches the edge of its knowledge. Leaders often assume that accuracy is the primary risk, but Naveh argues that the real exposure emerges when an agent encounters uncertainty, the moment where a traditional automation system would stall, fail silently, or produce a hallucinated output.

As he explains:

“Everybody’s measuring accuracy, but they’re missing the part of what happens when the agent doesn’t know. You don’t want it to stop, and you definitely don’t want it to hallucinate. You want the agent to raise its hand, reach out to a subject‑matter expert, and ask a specific series of questions to get itself out of the hole.”  

– Yoav Naveh, Co‑Founder and Co‑CEO at Reindeer

Naveh’s point is that regulated workflows cannot rely on deterministic automation because the work itself is not deterministic. AML reviews, onboarding checks, treasury operations, and vendor validations all contain edge cases that humans resolve through context, institutional memory, and policy interpretation. Agentic AI must be engineered to replicate the behavior of escalation, not the illusion of certainty.

In practice, this means designing agents that can identify ambiguity, articulate what they don’t understand, and initiate structured dialogue with the right human expert — not simply hand off the case or produce an incomplete answer.

Ajay reinforces this operational requirement from the risk perspective, noting that divergence rarely begins with a catastrophic error; it begins with a small misalignment between what the agent assumes and what the workflow requires. In his experience, regulated environments need oversight mechanisms that detect these early signals and pause execution before risk compounds downstream. Swamy’s emphasis is that exception‑aware escalation is not a fallback; it is a primary control that prevents minor uncertainty from becoming regulatory exposure.

Their insights outline a model for exception‑aware agent design that is fundamentally different from traditional automation. Institutions must build workflows where:

  • Agents can detect when a case deviates from known patterns or policy logic.
  • Escalation is structured, not ad hoc — with predefined questions, routing paths, and subject‑matter checkpoints.
  • Human feedback is captured as data, not lost in email threads or side conversations.
  • Exceptions serve as training signals that strengthen the agent’s future performance through a continuous learning loop, rather than as recurring points of failure.

Naveh argues that escalation alone is not enough. He describes a two-loop model in which the first loop enables agents to engage subject-matter experts when they encounter uncertainty. In contrast, a second learning loop aggregates those interactions and incorporates the resulting knowledge into future execution. The objective is not merely to resolve exceptions, but to continuously expand agent coverage, improve performance, and reduce recurring points of failure over time.

The combined message is that in regulated BFSI operations, the measure of a mature agent is not how it performs when everything goes according to plan, but how it behaves when the plan breaks. Naveh defines the behavior as an agent raising its hand, asking the right questions, and learning from the interaction. Swamy defines the stakes: without exception‑aware escalation, institutions cannot prevent divergence or defend decisions.

They describe this as the operational backbone of agentic AI in financial services: systems that know when they don’t know, and workflows designed to catch uncertainty before it becomes risk.

Strategic Operating Model for Sustainable Agent Deployment

The closing theme in the series shifts from workflow execution to long‑term ownership, the part of agentic AI that institutions consistently underestimate, according to Yoav and Ajay.

Yoav argues that the ease of building agents has created a structural blind spot inside enterprises. Teams can now assemble impressive prototypes in days, but those prototypes quickly become liabilities when no one is accountable for maintaining, governing, or evolving them over time.

This is where Naveh introduces the hinge point of the entire operating‑model conversation:

“It’s so easy to build today, everybody wants to build it, and nobody wants to maintain. People get swept up in how easy it is to build and they don’t think about how it’s going to work in the long run. If you don’t put someone in charge of building a strategy, you end up dependent on tools you can’t govern.”  

Yoav Naveh, Co‑Founder and Co‑CEO at Reindeer

Naveh’s warning is not about technical debt; it’s about agent debt. When institutions build agents without a strategy, they inherit autonomous systems that behave independently but cannot be supervised, updated, or retired without disruption. He stresses that leaders must decide up front which workflows warrant internal ownership, which require specialized platforms, and which should be delegated to vendors due to long‑term governance obligations.

Yoav Naveh argues that agent evolution should be managed with the same discipline applied to software releases, including controlled rollout, testing, and evaluation. Ajay complements this view by emphasizing ongoing supervision, divergence monitoring, and model-risk oversight.

Naveh describes this as a two-loop governance model. The first loop enables agents to engage subject-matter experts whenever they encounter uncertainty, collecting the information needed to resolve the case. The second loop aggregates those interactions and feedback over time, allowing agents to expand their knowledge, improve coverage, and self-heal as new situations emerge.

A sustainable operating model for agent deployment, as described by Naveh and Swamy, is less about tools and more about institutional discipline. Their combined guidance points to four structural commitments leaders must establish if they want agents to evolve safely over time:

  • Clear ownership for agent lifecycle management — including versioning, testing, approval, and retirement.
  • A build‑versus‑buy strategy grounded in competitive advantage — internal builds only where proprietary data or unique workflows justify the investment.
  • A maintenance plan that prevents agent drift — ensuring updates reflect policy changes, regulatory shifts, and operational feedback.
  • A governance layer that supervises agent learning — so improvements strengthen compliance rather than introduce risk.

The combined message is that the challenge in BFSI is no longer building agents but sustaining them.

Naveh defines the risk by showing how organizations that build without strategy end up dependent on tools they cannot govern. Swamy defines the requirement by insisting that agent evolution must be supervised with the same rigor applied to mission‑critical software. Together, they close the series with the principle that agentic AI becomes an enterprise asset when institutions design an operating model to support it long after the prototype phase ends.

Share article

Subscribe to updates

Subscribe to weekly email with our best articles Financial Services updates that have happened in the last week.

Recommended from Emerj