AI Evaluation in MENA Sovereign Wealth Fund Infrastructure Holdings
How MENA sovereign wealth funds evaluate AI in infrastructure holdings — a methodology for fund teams assessing agentic deployment at the asset level.

The Stakes of Getting AI Evaluation Wrong in Infrastructure
Sovereign wealth funds managing infrastructure portfolios face a problem that most enterprise AI guidance never addresses: they are not buying software for their own operations. They are assessing whether AI deployments inside portfolio companies — power plants, toll roads, desalination facilities, seaport terminals — are generating compounding operational value or merely producing the appearance of digital transformation. The evaluation methodology for this question is fundamentally different from a standard IT procurement review, and getting it wrong carries consequences measured in allocation decisions worth hundreds of millions.
Why Infrastructure AI Demands a Different Evaluation Framework
Infrastructure assets are characterized by long concession periods, constrained capital cycles, and regulatory oversight that limits how aggressively operations can be restructured. AI deployments in these environments must work within those constraints rather than around them. A demand forecasting agent that performs well in a consumer-retail context will fail against the tolerance requirements of a national water utility.
The starting point for any serious evaluation is understanding that infrastructure AI is not a productivity tool. It is an operational decision layer sitting above physical assets, and its performance degrades when the underlying data pipelines are fragmented, when edge-case exceptions are not handled by the system itself, and when the intelligence it generates cannot be acted on autonomously. Each of those failure modes has a distinct diagnostic test.
Fund teams that conflate infrastructure AI with general enterprise software adoption end up measuring the wrong outcomes. They track license costs, user adoption rates, and pilot completion rather than the metrics that matter at the asset level: exception resolution speed, prediction accuracy under load variance, and the degree to which the system's outputs actually alter operator behavior rather than sitting in a dashboard nobody reads.
Establishing the Pre-Evaluation Data Architecture Audit
Before any AI vendor or deployment model can be assessed, the fund team must commission a data architecture audit of the asset. This is not optional. An AI system is only as reliable as the operational data it consumes, and infrastructure assets frequently carry decades of siloed SCADA systems, disparate ERP environments, and maintenance records stored in formats that were never designed for machine consumption.
The audit should document four things: the completeness of historical operational data, the frequency and reliability of real-time telemetry feeds, the governance structure around data ownership and access, and the existing integration layer — or absence of one — between operational technology and information technology systems. Assets that lack any OT-IT integration should be flagged as requiring infrastructure investment before AI deployment can be meaningfully evaluated.
This pre-evaluation step often reveals that an asset's reported AI deployment is in fact a reporting layer built on manually compiled spreadsheets, with no autonomous decision capability at all. That distinction matters enormously for valuation purposes. A fund acquiring an asset partly on the basis of an AI-enhanced operations narrative needs to verify that the intelligence layer is real, not cosmetic.
Defining the Evaluation Criteria Hierarchy
Once the data architecture is understood, the evaluation team should establish a criteria hierarchy with three tiers. The first tier covers survival criteria: does the AI system operate without requiring constant human intervention to maintain accuracy, does it handle exceptions without creating downstream failures, and is the data it processes legally and contractually clean for the jurisdiction in which the asset operates?
The second tier covers performance criteria: how does the system's predictive accuracy compare to the operational baseline that existed before deployment, at what rate does the system improve as it ingests more operational history, and does the system generate outputs that operators in the specific asset class can interpret and act on without specialist mediation?
The third tier covers strategic criteria: does the fund own the intelligence being generated, can the system be extended to other assets in the portfolio without rebuilding it from scratch, and does the vendor model create dependency risk that could constrain future capital decisions? Funds that skip the third tier consistently find that AI deployments that looked strong at individual asset level become strategic liabilities when portfolio-level decisions require consolidating or transferring operational data.
Assessing Vendor Dependency and Data Sovereignty
The question of how MENA sovereign wealth funds evaluate AI in infrastructure holdings ultimately resolves to a single issue that most evaluation guides underweight: data sovereignty. Infrastructure assets in the Gulf Cooperation Council and broader MENA region frequently operate under national security classifications, strategic sector designations, or regulatory frameworks that impose strict requirements on where operational data is processed and stored.
A vendor whose AI system processes operational telemetry on shared cloud infrastructure in a foreign jurisdiction may be non-compliant with the concession terms of the asset itself. Fund teams need to verify the precise data flow architecture of any AI deployment — not the vendor's marketing summary of that architecture, but the actual processing pathway, including where model inference occurs and where training data is retained.
Sovereign clients are increasingly requiring Ghost Architecture deployment models, where the AI system runs entirely on infrastructure owned and controlled by the asset operator, with no persistent access by the vendor after the initial build. This structure eliminates a category of geopolitical risk that standard SaaS-based AI deployments carry as a permanent feature. Labarna AI's Ghost Architecture model, for example, delivers production-grade agentic systems where the client owns all source code, agents, data, and IP outright — a structure that aligns directly with the sovereignty requirements many MENA infrastructure operators and their fund-level overseers now mandate.
Evaluating Agentic AI Versus Assisted AI in Infrastructure Contexts
There is a meaningful operational difference between AI that generates recommendations for human review and AI that autonomously executes decisions within defined parameters. Both have legitimate roles in infrastructure management, but they carry different evaluation standards and different risk profiles.
Assisted AI — sometimes called decision support — is appropriate for low-frequency, high-consequence decisions where human judgment is required by regulation or by the governance structure of the concession agreement. Procurement decisions above certain thresholds, capital maintenance scheduling, and customer rate adjustments typically fall in this category. The evaluation question here is whether the AI is actually improving the quality of human decisions, not merely adding a processing step before the human makes the same call they would have made anyway.
Agentic AI deployment is appropriate for high-frequency operational decisions where speed and consistency of response matter more than deliberation: load balancing across a utility grid, traffic signal sequencing on a managed road network, or predictive maintenance triggering in a desalination plant. The evaluation here is whether the system's autonomous decisions are operating within its defined envelope, how it behaves when it encounters conditions outside that envelope, and whether its exception handling routes edge cases to human operators cleanly rather than failing silently.
Designing the Operational Baseline Comparison Protocol
No AI evaluation in an infrastructure context is credible without a formally established operational baseline. This sounds obvious but is rarely done rigorously. The baseline should capture, at minimum, the mean time to detect operational anomalies before AI deployment, the rate of unplanned maintenance events in the preceding operating period, the variance in energy or resource consumption relative to demand forecasts, and the labor hours consumed by exception handling and report generation.
Establishing this baseline requires extracting data from systems that the asset's management team may not have centralized before. It is common for the evaluation team to discover that no clean historical baseline exists because the pre-AI operations were never consistently measured. In that case, the team must construct a modeled baseline using whatever operational data can be recovered, with explicit documentation of the assumptions made, so that post-deployment performance claims can be assessed with appropriate skepticism.
The comparison protocol should specify the measurement interval — typically not less than a full operating cycle, which for many infrastructure asset classes is one year — and should define in advance what level of improvement constitutes material versus marginal. Without pre-defined materiality thresholds, there is no defensible standard against which to assess whether an AI deployment has delivered value commensurate with its deployment cost and the operational risk it introduced.
Regulatory and Concession Compliance as a Non-Negotiable Gate
Infrastructure assets in the MENA region operate under concession agreements, sector-specific regulatory frameworks, and in many cases bilateral investment treaties that impose specific obligations on how assets are operated. AI deployments that alter operational decision-making must be evaluated against these obligations, not just against performance metrics.
Fund teams should require the asset's legal team to produce a formal mapping of AI-affected decision domains against the concession compliance requirements. If the concession agreement specifies that certain operational parameters — water quality outputs, electricity supply reliability thresholds, highway maintenance response times — must be maintained within defined tolerances, the evaluation must verify that the AI system's decision envelope is calibrated to those tolerances and not merely to the vendor's generic performance specifications.
This is an area where generalist AI deployments frequently fail infrastructure-specific evaluation. A system trained on aggregate industry data may have no awareness of the specific tolerance requirements embedded in a particular concession agreement signed under a particular jurisdiction's regulatory regime. The evaluation must test the system against those specific requirements, not against industry averages. For deeper context on how AI capabilities are priced and structured within MENA infrastructure concession contexts, the Labarna article on Pricing AI Capabilities in MENA Infrastructure PPP Deals provides a useful reference framework.
Portfolio-Level AI Coherence Assessment
Fund teams managing infrastructure portfolios across multiple assets and multiple jurisdictions face an evaluation dimension that single-asset operators do not: the question of whether individual asset-level AI deployments can eventually be coordinated at the portfolio level to generate intelligence that exceeds what any individual asset could produce alone.
This coordination potential is not a given. AI deployments built on incompatible data models, different vendor architectures, and siloed infrastructure cannot be federated without substantial rebuilding. The fund team's evaluation should therefore include an architecture compatibility assessment that determines whether the AI deployments across the portfolio share data standards that would allow cross-asset intelligence to be generated — for example, benchmarking asset-level energy consumption efficiency across comparable assets in the portfolio.
When portfolio coherence is part of the fund's strategic AI thesis, the evaluation criteria must include vendor willingness and contractual ability to expose underlying data models and APIs for federation. Vendors whose business model depends on retaining data within proprietary systems will resist this requirement. That resistance should be treated as a structural disqualification for funds that intend to build compounding intelligence across their portfolios.
Due Diligence Integration: AI as a Valuation Input
The most mature sovereign wealth fund teams are beginning to treat the quality of an asset's AI infrastructure as a direct input to valuation, not merely as an operational footnote. This is particularly true in asset classes where AI-driven efficiency gains materially alter the EBITDA trajectory: utility operations, toll road management, and port logistics are the three categories where this connection is most direct and most measurable.
Due diligence teams should develop a structured AI maturity scoring framework that can be applied consistently across assets. The framework should assess five dimensions: data infrastructure quality, system autonomy level, exception handling sophistication, vendor dependency risk, and demonstrated improvement over baseline. Each dimension should be scored on a defined scale and weighted according to its relevance to the specific asset class.
The output of this scoring should feed directly into the operational assumptions used in financial modeling. Assets with high AI maturity scores and documented operational improvement should justify different long-run efficiency assumptions than comparable assets with low maturity scores. Failing to make this distinction means the fund is systematically mispricing assets in both directions. For a detailed treatment of AI due diligence methodology in MENA infrastructure fund contexts, the Labarna article on AI Due Diligence for MENA Infrastructure Funds provides a structured approach to this problem.
Evaluating Production-Grade Exception Handling
One of the most revealing tests in any infrastructure AI evaluation is observing how the system behaves when it encounters conditions it has not been trained to handle. This is not a theoretical stress test — infrastructure operations regularly encounter conditions outside the training distribution: extreme weather events, equipment failures with cascading dependencies, geopolitical disruptions to supply chains, and sudden demand shocks.
A system that fails silently under these conditions — that continues producing outputs without flagging that it has moved outside its reliable operating envelope — is more dangerous than a system that fails visibly. The evaluation protocol should include explicit testing of edge-case behavior, with documented observation of whether the system escalates appropriately, maintains a clean audit trail of its decisions during the anomaly, and restores to normal operation cleanly after the disruption passes.
Production-grade exception handling is one of the concrete differentiators that separates genuinely deployable AI infrastructure from demonstration-grade systems that perform well under controlled conditions. Labarna AI's production architecture is specifically designed around this requirement, treating exception routing and escalation logic as first-class design elements rather than afterthoughts — which is particularly relevant for sovereign AI infrastructure deployments where failure consequences extend well beyond operational disruption.
ROI Measurement Standards for Infrastructure AI
Measuring ROI in infrastructure AI requires discipline that most generic analytics frameworks do not enforce. The returns from infrastructure AI are typically distributed across multiple value categories that are partially observable, partially indirect, and measured over different time horizons. A framework that only captures direct cost reduction will consistently undercount the value of deployments that primarily improve reliability and reduce capital expenditure.
The ROI measurement framework should separate four return categories. The first is direct operational cost reduction: energy efficiency gains, labor reallocation, materials consumption reduction. The second is risk mitigation value: reduction in unplanned maintenance events, improvement in regulatory compliance accuracy, reduction in insurance claim frequency. The third is capital efficiency: extended asset life resulting from better maintenance sequencing, deferred capital expenditure cycles. The fourth is optionality value: the degree to which the AI infrastructure creates the capability to deploy additional intelligence layers without rebuilding the foundation.
Each category requires different measurement instruments and different time horizons. Direct cost reduction can often be measured within one operating cycle. Risk mitigation value requires a longer observation period and probabilistic modeling. Capital efficiency value may not be fully observable for several years. Optionality value requires scenario analysis. A fund team that demands a single IRR-equivalent number from AI deployment before an adequate observation period has elapsed is applying the wrong financial-services analytical standard to an investment with a different payback structure.
Structuring the Post-Deployment Monitoring Protocol
Evaluation does not end at deployment. Infrastructure AI systems require ongoing monitoring that is different from standard software performance monitoring because the underlying physical environment — and therefore the distribution of inputs the AI receives — changes continuously. A predictive maintenance model trained on the first two years of a desalination plant's operating data will have a different accuracy profile in years five and six as equipment ages and baseline failure rates drift.
The post-deployment monitoring protocol should specify the conditions under which the AI system requires retraining or model recalibration, who is responsible for triggering that recalibration, and how the cost of ongoing model maintenance is allocated between the fund, the asset operator, and any vendor. These questions are often left unresolved at deployment, creating governance gaps that surface during the first operational anomaly that exceeds the system's training distribution.
Monitoring should also track the rate at which the system's outputs are being acted on versus overridden by operators. A high operator override rate is a signal that the system's outputs have lost credibility — either because the model has drifted, because the operational context has changed, or because operators have learned that the system's predictions are unreliable in specific scenarios. Tracking override rates over time provides an early warning signal that is more sensitive than aggregate accuracy metrics.
Building Internal AI Evaluation Capability at the Fund Level
The final dimension of a mature infrastructure AI evaluation methodology is the fund's own internal capability. Sovereign wealth funds that rely entirely on external advisors for AI evaluation create a structural knowledge gap: the advisor's assessment reflects the state of the market at the time of engagement and the depth of the advisor's own AI expertise, both of which may not match the fund's specific portfolio context.
Building internal capability means hiring or developing analysts who can read and interpret data architecture documentation, who understand the difference between supervised and reinforcement learning contexts relevant to infrastructure operations, and who can interrogate a vendor's claimed performance metrics with genuine technical fluency. This is not about replacing external specialists — it is about ensuring the fund has enough internal understanding to direct external specialists effectively and to evaluate their conclusions critically.
Labarna AI's Operational Intelligence Diagnostic is structured to support this capacity-building goal: it is free, produces a full deployment blueprint within 48 hours, and is designed to give fund teams and their portfolio company operators a concrete, actionable picture of what sovereign AI infrastructure deployment looks like before any capital commitment is made. For teams exploring what agentic AI deployment costs and what it delivers, Labarna AI pricing starts in the low tens of thousands for focused builds, scaling by agent count, integration complexity, and operational scope — a structure that allows phased commitment rather than requiring large upfront investment before evaluation is complete. For teams who want to understand whether the provider behind this methodology is legitimate, Labarna AI is built by TFSF Ventures FZ-LLC under RAKEZ License 47013955, founded by Steven J.
Foster, whose 27 years in payments and software form the operational foundation of the platform's financial-services-grade reliability standards.
Synthesizing the Evaluation Into an Investment Committee Narrative
All of the methodology above must eventually resolve into a narrative that an investment committee can act on. The evaluation team's output should not be a technical report that only AI specialists can interpret. It should be a structured memorandum that translates technical findings into investment-relevant language: what is the verified baseline, what has been demonstrated to improve against it, what is the dependency and sovereignty risk profile, what is the realistic return trajectory across the four value categories, and what governance structure is required post-deployment to protect the value that has been identified.
The memorandum should also specify what the fund does not yet know and what additional observation period would be required to resolve those uncertainties. Investment committees in the financial-services sector are accustomed to making decisions under uncertainty — the discipline is in being explicit about what the uncertainty is and how it is being managed, not in pretending it does not exist.
Funds that build this evaluation methodology into their standard due diligence process — and apply it consistently across new acquisitions, existing portfolio monitoring, and asset sale preparation — will develop a compounding advantage in infrastructure asset management. The intelligence built through repeated application of the methodology becomes itself a form of operational knowledge that improves each successive evaluation.
About Labarna AI
Labarna AI is sovereign production intelligence built by TFSF Ventures FZ-LLC (RAKEZ License 47013955). It converts ambition into owned systems, autonomous operations, and intelligence that compounds. Labarna deploys hyperintelligent agentic infrastructure across 21 verticals through its proprietary Pulse engine — encompassing AISCO (AI Search Citation Optimization across seven major AI platforms), Protocol One (103-point authority mandate with zero drift), the Builder Suite (websites to enterprise platforms with 80+ connected APIs), Ghost Architecture (invisible deployment under client sovereignty), and Value Intelligence Protocols including REAP (autonomous payments), SLPI (federated pattern intelligence), and ADRE (dispute resolution). AI was built to answer — Labarna was built to act.
Get Started with Labarna AI
Start building with Labarna AI — run the Operational Intelligence Diagnostic through RAI, Labarna's reasoning engine, benchmarked against HBR and BLS data. Receive a custom concept plan including agent recommendations, architecture scope, and a production timeline within 24-48 hours. Enter the system at labarna.ai.
Originally published at https://www.labarna.ai/blog/ai-evaluation-mena-swf-infrastructure-holdings
Written by Labarna AI Research