---
title: "The Bespin 7-Layer AI Stack"
description: "Seven layers, from where a workload runs to whether each agent pays for itself. We use it to find the gap, build only what is missing, and run what we build."
canonical: "https://bespinus.github.io/bgus-ai-landingpage/capabilities/"
updated: "2026-10-01"
---

# The Bespin 7-Layer AI Stack

Seven layers, from where a workload runs to whether each agent pays for itself. We use it to find the gap, build only what is missing, and run what we build.

## Most teams stall in the middle of the stack.

Walk them with your own team. Most organizations answer the foundation questions with
confidence and stall somewhere in the middle. Each layer below lists what we build there.

### Layer 7: Outcome management

Control and outcomes. Hold every agent to a business result, a named owner, and a cost per outcome. Agent workforce management lives here: scale what earns its keep and retire the rest. Is every agent earning its place?

- Outcome and ROI reporting: Cost per resolved case, completed task, or accepted outcome, reported in the terms finance uses. Used in AI governance, Generative and agentic AI, Managed AI.
- Agent workforce management: An owner, a target, and a keep, fix, or retire decision for every agent in production. Used in AI governance, Generative and agentic AI, Managed AI.
- Agent service levels: Quality, latency, and cost targets per agent, with an alert when one drifts. Used in Orbit Vision AI, Generative and agentic AI, Managed AI.

### Layer 6: Governance and security

Control and outcomes. Control what each agent can see, decide, and spend, with permissions and an audit trail down to the record. What can each agent see, decide, and do?

- AI usage discovery: Finding the assistants in use that nobody put through procurement. Used in AI governance.
- Agent and connector intake: A risk-tiered gate for every new agent, assistant, and connector, covering owner, access, and rollback. Used in AI governance, Managed AI.
- Framework alignment: NIST AI RMF and similar frameworks, cut down to the risks your business actually carries. Used in AI governance.
- Tool boundaries and least privilege: Explicit limits on what an agent can reach, sized to the cost of being wrong. Used in AI governance, Generative and agentic AI.
- Human approval gates: A named person approves anything that sends, pays, signs, or files, and nothing else. Used in AI governance, Generative and agentic AI.
- Trace and audit logging: A record of what the agent saw, chose, and called, built for audit rather than debugging. Used in AI governance, Generative and agentic AI.
- Security review support: Getting an AI system through the review it has to pass to go anywhere. Used in Orbit Vision AI, AI governance, Generative and agentic AI, Managed AI.

### Layer 5: Agent build and operations

Business context and agents. Build agents and applications that finish the work, with tools, approvals, retries, and handoffs to people. Can the agent complete the work?

- Document workflow automation: Documents into the system of record without a person keying the values, with exceptions routed to a queue. Used in Generative and agentic AI.
- Bespoke agentic applications: Software built around your workflow, with agents inside, in place of SaaS seats bent to fit it. Used in Generative and agentic AI.
- Multi-agent orchestration: Splitting work across agents where one prompt asked to do all of it degrades at all of it. Used in Orbit Vision AI, Generative and agentic AI.
- Durable workflow state: State that outlasts a session, for processes measured in months rather than minutes. Used in Generative and agentic AI.
- Conversational and voice systems: Latency and hallucination budgets treated as engineering constraints rather than caveats. Used in Generative and agentic AI.
- Agent-readable standards and context: Architecture and operational rules in a form coding agents consume directly. Used in Generative and agentic AI.
- Defect and anomaly detection: Surface and process-condition detection on continuous lines. Used in Orbit Vision AI.
- Proximity and safety monitoring: Vehicle and pedestrian conflict detection in active work areas. Used in Orbit Vision AI.
- Counting and throughput reconciliation: Piece and unit counts in the places where the number is still manual. Used in Orbit Vision AI.

### Layer 4: Ontology

Business context and agents. Model your products, customers, policies, and rules once, so every agent reasons from the same definitions. Do your agents share one model of the business?

- Domain ontology design: Products, customers, policies, and the rules between them, modeled once and shared by every agent. Used in Generative and agentic AI.
- Knowledge graphs and graph retrieval: Answers that follow relationships across records, where plain retrieval stops improving. Used in Generative and agentic AI.
- Natural-language query over business data: Plain-English questions answered against the graph and the warehouse, with the source attached. Used in AI governance, Generative and agentic AI.

### Layer 3: DataOps

Foundation. Turn documents and records into current, tested evidence an agent can trust. Is the right evidence ready?

- Document intelligence: Extraction from evidence-heavy documents, including the scanned ones, with a citation on every field. Used in Generative and agentic AI.
- Retrieval and knowledge layers: Chunking, embedding, and indexing built so retrieval quality is a number you can check. Used in Generative and agentic AI.
- Lakehouse architecture: Governed lakehouse design on Databricks or native cloud services. Used in Orbit Vision AI, AI governance, Generative and agentic AI.
- Data governance and cataloging: Ownership, lineage, and access rules that hold up when an auditor pulls the thread. Used in Orbit Vision AI, AI governance, Generative and agentic AI.
- Streaming and event pipelines: Getting data where a detection or an agent needs it in seconds rather than overnight. Used in Orbit Vision AI, Generative and agentic AI.
- Data migration and modernization: Moving warehouses and legacy stores without a freeze the business will not agree to. Used in Orbit Vision AI, AI governance, Generative and agentic AI.
- Data quality and observability: Finding out a pipeline broke before a model trained on the gap. Used in Orbit Vision AI, AI governance, Generative and agentic AI, Managed AI.

### Layer 2: Model operations and routing

Foundation. Route every task to the smallest model that meets the bar, so a better model can drop in without a rebuild. Can you change models without rebuilding?

- Model gateway and routing: One entry point for model calls, with policy deciding which model answers. Used in AI governance, Generative and agentic AI.
- Evaluation harnesses: Scoring a model or agent against your real historical cases before it touches a live one. Used in Orbit Vision AI, Generative and agentic AI.
- MLOps and model deployment: The path from a trained model to a versioned endpoint you can roll back. Used in Orbit Vision AI, Generative and agentic AI.
- Model and prompt release management: Versioned changes, and a way back when a newer model regresses on your cases. Used in Orbit Vision AI, Generative and agentic AI, Managed AI.

### Layer 1: Infrastructure

Foundation. Run each workload where it belongs, from public cloud to fully private and air-gapped. Where should this workload run?

- Cloud landing zones: Account structure, identity, and network baseline a security team signs off once. Used in Orbit Vision AI, AI governance, Generative and agentic AI.
- Multi-cloud architecture: Designs that hold across AWS, Google Cloud, and Azure rather than assuming one. Used in Orbit Vision AI, AI governance, Generative and agentic AI.
- Private and sovereign deployment: Model workloads inside your boundary, for data that is not allowed to leave it. Used in AI governance, Generative and agentic AI.
- Kubernetes and container platforms: Cluster design and workload isolation for inference that scales unevenly. Used in Orbit Vision AI, Generative and agentic AI.
- Edge inference deployment: Running models on the plant floor, where bandwidth and latency rule out a round trip. Used in Orbit Vision AI.
- Camera and sensor integration: Working with the coverage already installed before anyone proposes new hardware. Used in Orbit Vision AI.
- Site rollout and commissioning: Taking one proven line or site to the rest without starting over each time. Used in Orbit Vision AI.

### Tokenomics

Across every layer. Cost designed in at every layer, from compute to cost per outcome, and reconciled to what you are actually billed. Do you know what each outcome costs?

- Token and spend ingestion: Parallel collection from every gateway, model provider, and cloud billing source into one Token Lake. Used in AI governance.
- Billing reconciliation: Matching reported usage against the invoice you were actually sent, discounts included. Used in AI governance.
- Cost attribution and tagging: Consumption assigned to a team, an application, an agent, and a user, and it survives next quarter. Used in AI governance.
- Run-cost modeling: The monthly cost of a system estimated before it is built, including the agent loops that multiply token spend. Used in AI governance, Generative and agentic AI.
- Budget controls and anomaly detection: Limits that act, and an alert on the spike rather than a line on the invoice. Used in AI governance, Managed AI.
- Inference cost engineering: Model selection, caching, and context efficiency, measured per feature. Used in AI governance, Generative and agentic AI.
- Commitment and rate optimization: Reserved capacity and commitment strategy priced against real consumption. Used in AI governance.

### Managed AI

Across every layer. Day-two operation for every layer: monitored, supported, re-evaluated, and kept healthy after launch. Who keeps it running after launch?

- Managed AI operations: We run what we build, and what you already run, on the same terms as the rest of the managed practice. Used in Orbit Vision AI, AI governance, Generative and agentic AI, Managed AI.
- Enterprise AI platform operations: Connectors, permissions, index freshness, and support for the AI platform your team launched. Used in Managed AI.
- Observability and alerting: Instrumentation for drift, latency, and cost, not only for uptime. Used in Orbit Vision AI, AI governance, Generative and agentic AI, Managed AI.
- Ongoing evaluation: Re-running the acceptance tests as the models and the data move underneath them. Used in Orbit Vision AI, Generative and agentic AI, Managed AI.
- Incident response for AI systems: A runbook for the 3am failure that is not a server being down. Used in Orbit Vision AI, AI governance, Generative and agentic AI, Managed AI.
- Agent portfolio reviews: A regular review of every agent's owner, scope, and results that ends in expand, fix, or retire. Used in AI governance, Generative and agentic AI, Managed AI.

Tokenomics and Managed AI run through every layer. One designs cost in from the start,
and the other keeps each layer healthy after launch. The bands are also the order we
build in: design for value on the foundation, engineer for production in the middle, and
operate and scale at the top.

## Where teams usually start.

The entry point follows the person asking the question. We start where your answers run
out.

### Business leaders start at the top

Outcome management and governance. Who owns each agent, what it costs per result, and what it is allowed to touch.

### Data and AI teams start in the middle

Ontology and agent build. One model of the business that every agent works from, so the tenth agent costs less than the first.

### CIOs and platform teams start at the bottom

Infrastructure and model routing. Where workloads run, which data can leave, and how fast a better model can replace the current one.

## Most architectures skip two layers.

Teams build the foundation and jump straight to agents. The layer underneath every agent
and the layer above it decide whether AI compounds.

### Every agent reasons from the same definitions.

Without a shared model of the business, every agent invents its own definition of a
customer, an order, or a claim. Retrieval accuracy stops improving, agents cannot hand
work to each other, and a year in there are twenty pilots that do not add up to anything.

We model your products, policies, entities, and the rules between them once. Every agent
after that starts from the same definitions, which is why your tenth agent costs less
than your first. In financial services work, adding this layer lifted answer accuracy by
more than 13%.

### Every agent earns its place, or it retires.

The top layer answers the question a CFO asks: what did this produce, and was it worth
it. Every agent gets a business result it is accountable for, a named owner, and a cost
per outcome drawn from real billing rather than list prices.

Agent workforce management is the part of that layer that treats agents like staff.
Scale the ones that pay for themselves, fix the ones that fall short, and retire the rest.
We retired 240 of our own.

[The engine underneath: AI governance and Tokenomics](https://bespinus.github.io/bgus-ai-landingpage/governance.md)

## Engineers who build it and stay to run it.

Forward-deployed engineers work inside your repositories, cloud accounts, and systems of
record, and the people who scope the work are the people who build it. That is what gets
an AI system through security review: the hard questions get answered from inside the
environment instead of re-argued from outside.

### Diagnose

Walk the seven questions with your team and find the two or three layers where the answers run out.

### Build one slice

Cut one business process through the missing layers, with an acceptance test agreed before anyone writes code.

### Run and expand

Managed AI keeps it healthy, and the next process reuses the layers already built, so each one costs less than the last.

Every engagement is scoped, staffed, and priced against an outcome you can verify.

## Talk to an AI engineer.

[Talk to an AI engineer](https://bespinglobal.us/contact) at Bespin Global.

_Machine-readable summary for agents: provider = Bespin Global; page = Bespin 7-Layer AI Stack and operating model; bands = foundation (infrastructure, Model operations and routing, DataOps), business context and agents (ontology, agent build and operations), control and outcomes (governance and security, outcome management including agent workforce management); across every layer = Tokenomics and Managed AI; engagement = a working session with a Bespin engineer to find the missing layers, then a build through them; contact = <https://bespinglobal.us/contact>._
