MODUS: Model Orchestration, Definitions & Usage. Secured.
The AI gateway for Azure: every department uses AI, and you see the costs per project.
MODUS brings all of your company's AI applications together in one place and runs on Kubernetes: in your own Azure tenant, on-premises or in another cloud, with no tie to a single cloud provider. You decide which model each application uses, how much it may consume and what gets stored.
Running in production at an international pharmaceutical group.
Cost per project
AI services
- Sales
- Service
- Quality
- Production
MODUS
on Kubernetes
No tie to a single cloud provider
- Azure
- On-premises
- another cloud
Azure
- Language models
- Machine learning
- Document analysis
- Search
Beyond Azure
- other providers
01Problem
Why AI usage is hard to control without a central point
Without a central point, model choice, system messages, API keys and logging live inside each application, and nobody sees AI usage as a whole. When several applications share one Azure OpenAI deployment, the Azure bill doesn't show which project used how many tokens. Switching models means changing code in every affected application. A pilot that was supposed to end in March keeps running until someone revokes the key.
At the same time, employees use AI outside the official route. 42 percent of German companies with 20 or more employees know or suspect that staff use generative AI such as ChatGPT through private accounts at work. 26 percent provide their own access, rising to 36 percent among companies with 100 to 499 employees (Bitkom, October 2025). This shadow AI grows wherever an official route is missing.
ISO 27001 requires rules for the use of cloud services (Annex A 5.23) and logging of relevant activities (Annex A 8.15). For external AI services, this means you need to be able to say which application uses which service under which rules. Without a central point, you can only answer that application by application.
02How MODUS works
How MODUS provides AI services from Azure and beyond
MODUS sits between your applications and the AI services and runs in your Azure tenant. Which services a project uses depends on its requirements, from machine learning models to language models. For language models, MODUS works as an Azure OpenAI gateway: applications don't call a model directly, they call an LLM definition that you maintain centrally.
LLM definitions
A definition sets which model is used, which system message it receives, which usage limits apply and how long the definition is valid. By default, language models come from Azure OpenAI. If you want to bypass Microsoft Foundry, MODUS also connects providers outside Azure, such as Anthropic, Mistral, the OpenAI API or services like OpenRouter.
Central API
For language models, your applications create completions through the MODUS API based on a definition. MODUS forwards the request to the model set in the definition.
Log, usage and chat sessions
Every AI request runs through MODUS and is recorded there. You decide what gets stored: completions and chats can be stored permanently, kept in memory only, or not stored at all. In the last case, only metadata such as token usage remains. When storing permanently, MODUS removes the link to user and session. MODUS assigns token usage to each project. MODUS also manages chat sessions centrally instead of each application handling them on its own.
Tools and company data
MODUS supports Retrieval Augmented Generation (RAG) and the Model Context Protocol (MCP). From Azure, you connect Azure AI Search and Azure Document Intelligence: Document Intelligence extracts text and tables from PDFs, scans and forms, and Azure AI Search finds the relevant passages so a model can answer based on your documents.
A run for a language model
- You create an LLM definition: model, system message, usage limits, validity.
- An application requests a completion for this definition through the MODUS API.
- MODUS checks validity and usage limits, forwards the request to the model and records it in the log.
03Result
What changes for IT management and development
Model choice, system messages, usage limits and storage rules sit in one place, and every AI request can be traced there.
| Without a central point | With MODUS | |
|---|---|---|
| Switching models | code change in every application | change in the LLM definition |
| System messages | in the code of each application | central in the definition |
| Usage limits | quota per Azure OpenAI deployment, not per application | per definition |
| Usage per project | not visible when applications share a deployment | token usage per project |
| End of a pilot | access continues until someone revokes it | definition expires at the set date |
| Storage of requests | different in each application | set centrally: permanent, in memory only, or metadata only |
| Traceability | logs, as far as each application writes them | one central log of all AI requests |
| RAG and MCP | built separately in each application | central in MODUS |
If you offer your employees your own AI applications as an alternative to private accounts, they all run through the same approved route.
MODUS runs in production in the Azure tenant of an international pharmaceutical group, where it provides the AI models centrally to all projects and assigns token usage to each project.
You can run MODUS yourself or hand over operations to us as part of ASTRA.
04Fit
Who MODUS is for
MODUS fits if
- several projects use AI services from Azure or are about to,
- your IT should decide which application uses which model and to what extent,
- you don't want to build your own API management for it,
- you run AI pilots with a fixed end date,
- your data protection officer should have a say in whether the content of AI requests is stored,
- you need to show for ISO 27001 or TISAX which applications use which AI services,
- your team shouldn't rebuild RAG and MCP in every application,
- you want to assign token usage to individual projects.
05Pricing
How much does MODUS cost?
Setting MODUS up in your Azure tenant costs EUR 7,500 once. After that, the licence depends on how many projects use MODUS. All prices excluding VAT.
Setup, one-time
EUR 7,500
Setting MODUS up in your Azure subscription, with integration, test and production environments, and briefing your team. Takes about one month.
up to 5 projects
EUR 1,499
/ month
up to 15 projects
Most common choiceEUR 2,499
/ month
unlimited
EUR 3,999
/ month
The licence covers the use of MODUS as well as feature and security updates. It runs for one year and renews for another year unless you cancel three months before the end of the term. MODUS runs in your own Azure tenant, and your data stays under your control. Azure consumption in your tenant, for example for models and infrastructure, is not included. You can run MODUS yourself or hand over operations to us as part of ASTRA. The number of users and applications does not affect the price. For source code handover and for group-wide use across several tenants, we put together an offer.
MODUS is currently available to early adopters who request it directly from us. Availability through Microsoft Marketplace is planned from 2027.
06FAQ
Frequently asked questions about MODUS
What is an AI gateway?
An AI gateway is a central layer between applications and AI models. Applications don't talk to a provider directly but to the gateway, which controls access, usage and logging in one place. MODUS is an AI gateway for Azure, runs in your own Azure tenant and also manages LLM definitions, chat sessions and the connection to RAG.
What is an LLM definition?
An LLM definition is the entry in MODUS through which an application uses a language model. It contains the model, system message, usage limits and validity period. When you change a definition, the change applies to every application that uses it.
We use Microsoft Foundry. Why do we need MODUS?
Microsoft Foundry (formerly Azure AI Foundry) provides the models and AI services. MODUS controls how your projects use them: which services a project gets, within which usage limits, and what usage it generates. On top of that come LLM definitions with system message and validity period, chat sessions and the connection to Azure AI Search and Azure Document Intelligence.
How does MODUS differ from the AI gateway in Azure API Management?
The AI gateway in Azure API Management requires an API Management instance, which many companies would have to set up first. MODUS doesn't need API Management and can be placed behind it later, once you introduce one. MODUS also gives finer control: an LLM definition holds the system message and validity period in addition to model and usage limits, and usage is assigned to each project. And because MODUS runs on Kubernetes, it is not tied to Azure.
Which AI services and providers does MODUS support?
MODUS provides all AI services available in Azure, from machine learning models to language models from Azure OpenAI. Which of them a project uses depends on its requirements. If you want to bypass Microsoft Foundry, MODUS also connects providers outside Azure, such as Anthropic, Mistral, the OpenAI API or services like OpenRouter.
What counts as a project?
A project is an application or an initiative with its own cost assignment in MODUS. If a department uses the same cost assignment for several applications, that is one project. The number of users and applications behind it does not affect the price.
What happens if we don't renew the licence?
MODUS is deactivated at the end of the term. The MODUS database with your definitions and settings is in your subscription and stays there.
Can MODUS assign AI usage to individual projects?
Yes. MODUS records the token usage of every request and assigns it to the respective project. This shows which project uses how many tokens, even when several applications share one Azure OpenAI deployment. That is the basis for charging AI costs internally.
What does MODUS store from your AI requests?
You decide. Completions and chats can be stored permanently, kept in memory only, or not stored at all. In the last case, only metadata such as token usage remains. When storing permanently, MODUS removes the link to user and session, while the chat content itself stays unchanged. If no content should be kept permanently, choose in memory only or metadata only.
Where does MODUS run?
By default, in an Azure subscription in your own tenant, on Azure Kubernetes Service with separate integration, test and production environments. Definitions and settings are stored in a MODUS database in the same subscription. Because MODUS runs on Kubernetes, it can also run on-premises or in another cloud. You can run MODUS yourself or hand over operations to us as part of ASTRA.
What do we need for the setup?
An Azure subscription for us to deploy MODUS into, and your requirements for domain and certificates. We handle the entire setup with Bicep, and it takes about one month.
How does MODUS use Azure AI Search and Azure Document Intelligence?
Azure Document Intelligence extracts text, tables and structure from PDFs, scans and forms. Azure AI Search searches that content. MODUS connects both services so a model can answer questions based on your own documents, such as maintenance manuals or inspection reports. This approach is called Retrieval Augmented Generation (RAG).
Is MODUS a Microsoft product?
No. MODUS is a product of DEVDEER GmbH. DEVDEER is a Microsoft Cloud Solution Provider.
Next step
Ready to create impact?
Tell us briefly what it’s about – by email or in a non-binding conversation. We listen, ask the right questions, and show how we can help in a solution-oriented and pragmatic way.
What happens next
- 01
Describe your case
Three fields, no sign-up. Two minutes is enough.
- 02
Personal reply
Stefanie Heine gets back to you within one business day.
- 03
Non-binding first conversation
We listen, ask the right questions, and show how we can help.
- hello@devdeer.com
- +49 (0) 391 - 55 68 00 5 0
- Herderstraße 31, 39108 Magdeburg
Your contact

Stefanie Heine
Executive Assistant
0/500 characters
We respond within one business day.