Back to AI in Action
Case Study 04 · In-Flight Discovery Phase

Multi-Agent
Project Intelligence

Connecting risk, planning, commercial, BIM and IM insight — through bounded agents, an orchestrator, and approval-gated write-back.

Project leadership lacks a fast, consolidated view of project health. Risk, planning, commercial and BIM/IM information sit in separate documents and systems, requiring manual synthesis. A multi-agent approach is being designed — domain-specific agents coordinated by an orchestrator — to compress portfolio-grade reporting from days to seconds without compromising the audit trail. Pilot target: Grimsby to Walpole.

Azure OpenAI Copilot Studio Power Automate SharePoint Power BI
5 domain agents
Risk · Planning · Commercial · BIM · IM
< 90 sec
Target consolidated report time
Pilot target
Grimsby to Walpole programme
Approval-gated
Every write-back routes through Teams
01 · The 30-Second Story

Why This Exists

Large projects generate information across risk registers, schedules, commercial trackers, BIM models, IM systems and meeting records. Project leaders need joined-up answers — but the evidence sits in different places. The pilot explores how project teams can ask cross-domain questions and get a source-cited synthesis for human review, without a single AI ever crossing accountability boundaries.

Team
Darshan Ruikar · Mesut Pala · Ashish Ranjan
Project context
Large programme delivery; Grimsby to Walpole as the pilot target. Currently in discovery.
Problem
Project data is fragmented; reporting is slow; leadership questions cut across domains; manual synthesis is inconsistent.
Users
Project managers · PMO leads · delivery directors · commercial managers · BIM/IM leads.
AI role
Domain-specific analysis per agent · orchestrated synthesis · proposed actions for human approval · citation of source evidence.
Non-AI role
Source-of-truth registers · approval flows in Teams · write-back execution · audit-log management · permission control.
Main value (target)
Sub-90-second portfolio responses · cross-domain visibility · governed write-back · earlier surfacing of connected pressures.
Evidence
Architecture and governance pattern designed; pilot scope being confirmed. Discovery Targeted · pilot
Status
Discovery phase. No production deployment yet. Phase 1 (single-agent pilot) targeted at Grimsby to Walpole.
Reusable pattern
Bounded multi-agent orchestration with approval-gated write-back, citations, and per-domain accountability.
02 · Project Context

Why This Matters at Programme Scale

On a single project, manual stitching of cross-domain status is tolerable. Across a programme like Grimsby to Walpole — multiple disciplines, many active workstreams, fast-moving commercial and BIM data — the manual approach doesn't scale.

Lifecycle point
How multi-agent intelligence helps
Weekly status
On-demand consolidated view instead of multi-day stitch from siloed trackers.
Cross-domain triage
Patterns linking risk, schedule and commercial pressure surface early — not post-hoc.
Stage-gate evidence
Source-cited synthesis becomes part of the audit trail leading into review gates.
Leadership questions
Ad-hoc cross-domain questions answerable in under 90 seconds — not by a 2-day analyst run.
Lessons capture
Findings, evidence and approvals all land in the same audit trail — reusable for future projects.
Why this is a delivery story, not a tools story
The brief: keep the precision of human judgement and the audit trail of formal reporting, but compress the time between "data exists" and "leadership can act on it" from days to seconds.
03 · The Problem

Project Intelligence Lags Project Reality

By the time portfolio status reaches leadership, the underlying picture has already moved. The cost is decisions made on stale evidence. Three problems compound.

Problem 1 — Fragmentation

Data spread across SharePoint folders, Excel registers, BIM models, commercial trackers. No consolidated view; manual stitching every cycle.

Problem 2 — Latency

Status reports take days. Each weekly cycle compresses into a single update — late, incomplete, and stale before circulation.

Problem 3 — Silos

Risk, planning, commercial, BIM and IM rarely cross-check in real time. Patterns surface post-hoc, after consequences have landed.

The real question
How do we generate cross-domain project intelligence faster than days, without weakening the audit trail or merging accountabilities between disciplines?
04 · Why Multi-Agent Was Appropriate

A Single AI Can't Carry the Audit Trail

Plenty of tools can chat with one document or summarise one feed. The shape this problem needs is different — multiple bounded experts, each accountable for its own domain, coordinated without merging the accountabilities.

A single AI can
  • Summarise one document or feed
  • Answer questions on a single source
  • Draft text or run quick analysis
A single AI cannot
  • Hold deep expertise across 5 disciplines simultaneously
  • Maintain bounded data scope per domain (BIM ≠ commercial)
  • Keep approval flows separate per discipline owner
  • Scale to portfolio-grade analysis without losing precision
  • Produce an audit trail that maps a finding back to a single accountable agent
The shape multi-agent gives us
Bounded scope per agent, bounded data per agent, bounded approval per agent — coordinated by an orchestrator that consolidates without merging. The risk lead still owns risk; the commercial lead still owns commercial. AI carries the synthesis, humans carry the decisions.
05 · What Was Designed

Five Bounded Agents, One Orchestrator

Each agent reads only its discipline's sources, applies its own ruleset, and produces findings against its own approval flow. The orchestrator coordinates — it doesn't override. Discovery-phase design; not yet deployed.

Risk Agent

Reads the risk register and project documents. Proposes new risks; flags emerging exposure across workstreams.

Planning Agent

Tracks schedule variance, critical-path flags, and projected completion against the baseline.

Commercial Agent

Watches budget vs spend, CPI, and contingency drawdown. Surfaces commercial trends, not just point figures.

BIM + IM Agents

Drawing register currency, RFI compliance, clash tracking, information-management discipline against ISO 19650.

Coordinated by
An orchestrator agent routes user questions to the relevant domain agents, consolidates their findings into a single response, and routes any proposed write-back through the discipline owner's approval flow before it touches a system of record.
Workflow

How a question flows through the system

Trigger
Scheduled portfolio check, weekly report cycle, or a user question — "What's the top risk this week across the programme?"
A
Orchestrator routes The orchestrator parses the request, identifies which domain agents to consult, and dispatches scoped queries — each agent gets only its own data context.
B
Agents analyse in parallel Each domain agent reads its sources (SharePoint, registers, BIM models, commercial trackers), applies its ruleset, and returns findings with citations to underlying evidence.
C
Orchestrator consolidates Findings synthesised into a single coherent response — but each finding stays attributed to its source agent, preserving accountability.
D
Human approval before write-back If the response includes a proposed action (a new risk, a flagged variance), it routes through Teams to the discipline owner. Nothing reaches a system of record without explicit human sign-off.
Output
Consolidated portfolio response · Power BI dashboards updated · approved write-backs into SharePoint and registers · full audit trail per agent
06 · How the Process Changes (Intended)

From Days to Seconds — Without Losing the Audit

These describe the intended end-state — to be validated through the Grimsby to Walpole pilot. Items below are targets, not measured results.

Today · Manual
Target · Multi-Agent
Portfolio status: days to compile, weekly cadence.
On-demand, <90 second response.
Cross-discipline analysis: manual stitching across silos.
Orchestrated, parallel, attributed to source agent.
Pattern detection: post-hoc, after the consequence.
Continuous, agent-driven.
Write-back to systems: manual edits to registers.
AI-proposed, human-approved via Teams.
Audit trail: email + meeting minutes.
Per-agent, per-finding, per-approval, end-to-end.

Who would feel the change

Project manager

Faster cross-domain answers when leadership asks. Less time chasing trackers; more time deciding.

PMO lead

Continuous portfolio synthesis instead of weekly stitching. Pattern visibility days earlier.

Discipline lead

Findings stay attributed to their discipline; approval to update registers stays with the owner.

Delivery director

Fewer late surprises in cross-domain pressure points. Earlier visibility of risk × programme × commercial overlap.

Programme sponsor / client

Traceable audit trail per finding and per approval — defensible at stage gates and assurance reviews.

Bid / account lead

A responsible-AI PMO and owners-engineer proposition — bounded scope, citations, approval gates baked in.

Status — read this carefully
The system is in discovery. The architecture, agent boundaries, governance pattern, and approval flows are designed; no production deployment yet. First pilot is targeted at the Grimsby to Walpole programme — see §12 for the phased rollout plan.
07 · Value (Targeted)

Six Facets — Mostly Targeted, Not Yet Proven

Because this case study is in-flight, the value claims here are explicitly targeted rather than measured. Two facets — governance pattern and approval-gating — are confirmed by the design itself.

Time Targeted · pilot

Sub-90-second consolidated portfolio response. Today's manual stitching takes days. To be validated on Grimsby to Walpole.

Quality Targeted

Cross-discipline synthesis with citations to source evidence — earlier surfacing of connected pressures than manual review can produce. Validation through pilot.

Cost TBC

Reduced manual analyst time across reporting cycles, balanced against agent infrastructure and approval-flow latency. Quantified saving requires pilot evidence.

Risk Proven · pattern

100% of AI write-backs gated by human approval. Bounded per-agent scope means no merging of accountabilities. The risk-control pattern is designed-in, not bolted-on.

Governance Proven · pattern

Hosted in Arup; per-agent boundaries; citations on every claim; full audit trail per agent, per finding, per approval. The pattern itself is the value here.

Commercial Targeted · being productised

Foundation for a PMO / owners-engineer client offer combining responsible-AI agents with traceable governance. Productisation is part of the pilot work.

08 · Governance & Responsible Use

Why It's Safe to Pilot on a Live Programme

For an in-flight system, governance is the case for permission to deploy. The audit story has to be intact before the first write-back goes live.

Hosted inside Arup

Azure OpenAI running in Arup's secured environment. Programme data does not leave the controlled boundary.

AI proposes, humans decide

Every proposed action — new risk, flagged variance, register update — routes through Teams to the discipline owner before it touches a system of record.

Per-agent boundaries

Each agent reads only its discipline's sources. No agent sees the full programme — accountability stays attached to the right discipline lead.

Full audit trail

Every finding carries the agent that produced it, the source it was drawn from, and the human who approved any resulting action.

Citations on every claim

Findings link back to the source document, register row, or model element they were drawn from. No untraceable assertions.

Permission inheritance

Agent access respects existing project permissions — no agent can see what its human user couldn't already see.

For client conversations
The reason this pattern can be deployed on a live infrastructure programme is the same reason it generalises: each agent is bounded, every write-back is gated, and every finding is traceable to its source and its approver. Multi-agent doesn't mean unsupervised — it means parallelised under stricter governance, not weaker.
09 · The Technical Bit

For Specialists — Components, Validation, and What to Confirm

A deeper-dive section for technical readers. Discovery is still ongoing, so several architecture choices are open. The "to confirm" list is more important here than for the live cases.

For technical readers

"Bounded domain agents coordinated by an orchestrator, with evidence and human approval. AI synthesises; humans decide."

Components

Component
Purpose
Data connectors
Connect to SharePoint, Excel, schedules, risk registers, BIM/IM trackers and other approved project sources.
Knowledge index
Make project documents searchable and retrievable.
Agent instructions
Define each agent's role, boundaries, sources and output format.
Orchestration layer
Route questions to relevant agents and combine findings.
Evidence / citation layer
Show where claims came from — source document, register row, model element.
Approval workflow
Require human review before use or write-back.
Audit log
Record question, sources used, output and approval decision.

Validation activities (pilot)

Activity
What to test
Retrieval accuracy
Does each agent find the correct evidence?
Domain accuracy
Are findings technically correct for risk, planning, commercial, BIM and IM?
Cross-domain synthesis
Does the orchestrator combine findings without losing nuance?
Citation quality
Are claims traceable to source documents?
Permission control
Does the system respect project access rights?
Write-back safety
Are changes blocked until approved?
Hallucination control
Does the system say when evidence is missing rather than invent?
User acceptance
Do PMs and domain leads trust the outputs?
To confirm with pilot team — before publication
  • Platform: Copilot Studio vs Azure AI Studio vs custom Azure OpenAI stack — to describe the architecture accurately.
  • Which data sources are in scope — to define the evidence base.
  • Whether emails are included — to address confidentiality and data-sprawl constraints.
  • How permissions are inherited — to prove access control.
  • Whether retrieval-augmented generation (RAG) is used — to explain grounding.
  • Whether outputs are source-cited — to support trust.
  • Whether write-back is in scope and which systems are eligible — to define risk and governance.
  • Audit-log schema and retention period.
  • Current pilot status, target metrics, and validation plan — to avoid overclaiming.
10 · What Colleagues Can Reuse

Six Patterns to Lift — Even Before Pilot

Discovery has already produced reusable governance and architecture patterns that other teams can lift today. Adoption doesn't have to wait for the pilot.

Domain-bounded agent

One agent per discipline. No agent sees outside its scope. Apply to risk, planning, commercial, BIM, IM, or any other discipline.

Orchestrator pattern

Combine multiple domain findings into one answer — without merging accountabilities. Routing + consolidation + attribution.

Approval-gated write-back

No AI proposal touches a system of record without explicit human sign-off. Prevents uncontrolled updates to project systems.

Source-cited outputs

Every finding links back to its source. Improves trust and auditability — and makes the next reviewer's job possible.

Cross-domain prompt library

Standard prompts for programme leadership questions — "where are we behind?", "what's emerging?", "what's the critical path?"

Governance checklist

Per-agent data scope · approval flow · audit log spec · permission inheritance — use before piloting agents on client projects.

Reusable from this discovery — already
Agent-pattern libraryBounded-scope, approval-gated, citation-required.
Governance templatePer-agent data scope, approval flow, audit trail spec.
Orchestration modelRouting + consolidation + attribution without merging accountabilities.
11 · Client-Facing Proposition

How This Becomes an Advisory Offer

The multi-agent pattern underpins a productisable PMO and owners-engineer offer — bounded, governed AI agents that connect risk, planning, commercial, BIM and IM evidence with full audit trail and human-approved write-back.

Proposition statement
Arup helps clients deploy bounded, governed AI agents into PMO and owners-engineer functions — connecting risk, planning, commercial, BIM and IM evidence with full audit trail, citations on every claim, and human-approved write-back.

Offer components

Where this lands in the playbook
Maps to Service Line 01 — Business and Investor Advisory (BIA) for cross-domain evidence synthesis, and Service Line 03 — P&PM for governed PMO and owners-engineer intelligence.
12 · What's Next — Pilot Roadmap

From Discovery to Portfolio Scale

Four phases — currently in Phase 0. Each phase earns the right to the next; nothing scales until the audit pattern holds at the previous level. Outstanding "what to confirm" items are listed in §9.

Phase 0 · Now

Discovery

Confirming data formats, licences, and SharePoint scope. Designing per-agent boundaries and the approval flows that will make the pilot defensible.

Phase 1 · Pilot

Risk Agent · Grimsby to Walpole

Single-agent pilot. Document analysis, AI-proposed risks, human approval, full audit trail. Trust earned through transparent first runs.

Phase 2 · Scale

Full Portfolio

All five agents running across live projects, with Power BI dashboards consolidating the portfolio view. Orchestration patterns hardened.

Phase 3 · Advanced

Extend the Surface

PDF reading, semantic search across the corpus, mobile capture from site, email integration. Each capability inherits the existing governance baseline.

Earning the right to scale
The phased shape is deliberate. Discovery → single-agent pilot → multi-agent scale → extended surface. Each step has to demonstrate that the audit pattern holds before the next one starts. No leap-frogging — that's how an in-flight pilot turns into a credible client-grade capability.
Sponsored by Darshan Ruikar Work by Mesut Pala Pilot target Grimsby to Walpole programme
Final Takeaway

AI proposes, humans decide. Every item goes through Teams approval before write-back. Full audit trail maintained.

Bounded scope + Approval gates + Audit trail

The shape that lets an in-flight pilot be defensible on a live infrastructure programme is the same shape that lets it scale.

Continue