Tag: AI Gateway

Configure and deploy AI Gateway in Azure API Management for Microsoft Foundry (SC-500 Exam Prep)

This post is a part of the "SC-500: Implementing End-to-End Security Controls for Cloud and AI Workloads" Exam Prep Hub.
This topic falls under these sections:
Secure compute (20–25%)
   --> Implement security for AI
      --> Configure and deploy AI Gateway in Azure API Management for Microsoft Foundry


Note that there are 10 practice questions (with answers) at the end of each section to help you solidify your knowledge of the material. Also, there are 4 practice tests with 30 questions each available from the hub's main page below the exam topics section.

Overview

Microsoft Foundry provides services for developing, deploying, and operating generative AI applications, models, and agents. As organizations adopt AI at scale, they need a controlled way to manage access to models and AI tools.

An AI Gateway uses Azure API Management to provide a governed entry point between applications or agents and AI backends. It can centralize authentication, authorization, traffic control, token usage, monitoring, routing, and security policies.

For SC-500, the important concept is that the AI Gateway is not simply another model endpoint. It is a security and governance layer placed between AI consumers and the services they access.

Exam focus: Understand how to connect Microsoft Foundry to Azure API Management, configure the gateway, import models or tools, apply policies, and verify that traffic is actually being mediated by the gateway.


What Is an AI Gateway?

An AI Gateway is a set of capabilities in Azure API Management that helps organizations manage AI-related backends.

These backends may include:

  • Models deployed in Microsoft Foundry.
  • Azure OpenAI deployments.
  • Other supported model providers.
  • OpenAI-compatible model endpoints.
  • Remote Model Context Protocol (MCP) servers.
  • Agent-to-agent APIs.
  • Custom AI services.
  • Self-hosted models and endpoints.

The AI Gateway extends the existing API Management gateway. It is not a completely separate gateway product. Existing API Management capabilities, including policies, authentication, routing, monitoring, and networking, are used to govern AI traffic.

Why use an AI Gateway?

Without a gateway, each application or agent may connect directly to an AI model or tool. This can lead to:

  • Duplicated authentication logic.
  • Inconsistent security policies.
  • Uncontrolled model consumption.
  • Difficulty enforcing quotas.
  • Limited visibility into usage.
  • Excessive exposure of backend endpoints.
  • Different teams implementing different controls.
  • Difficulty changing model providers.

An AI Gateway provides a centralized control point for these concerns.

For example, several applications might use different model deployments, but all requests can pass through API Management where the organization applies:

  • Authentication.
  • Authorization.
  • Rate limits.
  • Token quotas.
  • IP restrictions.
  • Content safety policies.
  • Request and response transformations.
  • Logging and metrics.
  • Backend routing.
  • Load balancing.
  • Caching, where appropriate.

AI Gateway Architecture

A typical architecture contains the following components:

  1. AI consumer
    • Application.
    • Copilot.
    • Agent.
    • Development tool.
    • Automated workload.
  2. Microsoft Foundry resource or project
    • Hosts or manages model deployments, agents, and tools.
  3. Azure API Management instance
    • Acts as the AI Gateway.
    • Receives requests from consumers.
    • Applies policies.
    • Routes requests to the appropriate backend.
  4. AI backend
    • Microsoft Foundry model.
    • Azure OpenAI deployment.
    • External model provider.
    • MCP server.
    • Other supported AI endpoint.
  5. Monitoring and governance services
    • API Management logs and metrics.
    • Application Insights, where configured.
    • Microsoft Foundry telemetry.
    • Security monitoring and auditing.

The gateway sits between the client and the AI backend. This allows the organization to enforce common controls without requiring every client application to implement those controls independently.


AI Gateway in Microsoft Foundry

Microsoft Foundry can be integrated with an Azure API Management instance as an AI Gateway.

This integration allows organizations to govern AI resources from within the Foundry environment while retaining access to the more advanced configuration capabilities of Azure API Management.

Depending on the supported feature and configuration, the gateway can help govern:

Models

The gateway can provide:

  • Token quotas.
  • Rate limits.
  • Authentication.
  • Routing.
  • Usage monitoring.
  • Model access control.
  • Centralized governance across model deployments.

When AI Gateway is used with Foundry, model requests can be routed through the associated API Management instance. Model limits can be configured at the project level, helping prevent one project or team from consuming all available capacity.

Agents

Agents can be registered and governed through Microsoft Foundry. Governance can include:

  • Centralized inventory.
  • Traffic policies.
  • Throttling.
  • Content safety controls.
  • Monitoring.
  • Access management.

The exact capabilities depend on the agent type and the integration being used.

Tools

MCP tools can be routed through an AI Gateway so that requests pass through a controlled endpoint.

Policies can be applied to MCP traffic, including:

  • Authentication.
  • Rate limiting.
  • IP filtering.
  • Correlation IDs.
  • Logging and metrics.
  • Routing controls.

However, the Foundry MCP gateway integration has limitations. For example, only eligible MCP tools created after the gateway is connected may be routed through the gateway. Existing tools are not automatically changed to use the gateway.


Prerequisites

Before configuring an AI Gateway for Microsoft Foundry, verify the following.

Azure API Management instance

You need an Azure API Management instance that meets the requirements for the selected integration.

Supported service tiers and networking requirements can vary depending on:

  • Whether the gateway is public or private.
  • Whether the Foundry resource has public network access disabled.
  • Whether private endpoints are required.
  • Whether advanced networking is needed.
  • Whether the organization is using the dedicated AI Gateway tier preview.

The dedicated AI Gateway tier is a public preview feature. Preview capabilities, supported regions, limits, and service behavior may change. Organizations should validate preview features carefully before using them for critical production workloads.

Required permissions

The administrator configuring the integration generally needs permission to manage the API Management instance.

For some Foundry gateway scenarios, the required role is:

  • API Management Service Contributor, or
  • Owner

The exact permissions depend on whether the administrator is connecting an existing gateway, creating an instance, importing models, or managing policies.

Networking

If the Foundry resource has public network access disabled, the API Management instance must also be able to access the private Foundry resource.

Depending on the architecture, this may require:

  • A private endpoint.
  • A supported API Management tier.
  • Virtual network integration or injection.
  • Appropriate private DNS configuration.
  • Network rules that permit the required traffic.

A gateway cannot securely mediate traffic to a private backend if the gateway itself cannot reach that backend.

Backend access

The administrator must have access to the model, deployment, or tool backend being added.

For managed identity authentication, the managed identity must also have the required permissions on the backend resource.


Creating or Associating an AI Gateway

The exact portal experience can change, but the general process is:

  1. Sign in to Microsoft Foundry.
  2. Open the appropriate Foundry administration or resource configuration area.
  3. Open the AI Gateway configuration.
  4. Select Add AI Gateway.
  5. Select the Foundry resource to associate with the gateway.
  6. Select an existing API Management instance or create one if supported.
  7. Confirm the required permissions and networking configuration.
  8. Save the association.
  9. Add or import models, agents, or eligible tools.
  10. Configure API Management policies.
  11. Test requests through the gateway.
  12. Verify telemetry and policy enforcement.

The Microsoft Foundry portal provides an integrated configuration experience, while advanced policies and networking settings are managed in Azure API Management.


Importing Models into the AI Gateway

Azure API Management can import models from supported providers, including Microsoft Foundry and Azure OpenAI.

When importing a model, the administrator typically configures:

  • The model provider.
  • The backend endpoint.
  • The model or deployment name.
  • The API format.
  • Authentication.
  • Required headers.
  • Backend routing.
  • Policies.
  • Monitoring settings.

For Microsoft Foundry deployments, the import wizard can discover deployments automatically in supported scenarios.

OpenAI-compatible APIs

Many applications are designed to use the OpenAI API format. API Management can expose supported backends through OpenAI-compatible routes.

For example, an OpenAI-compatible model endpoint may use a route similar to:

/default/models/openai/v1/chat/completions

The exact gateway URL and route depend on the API Management configuration and the provider API format.

The important exam concept is that the client can use a consistent gateway endpoint while API Management handles the connection to the underlying model backend.


Authentication Options

Authentication must be configured separately for:

  1. The client calling the gateway.
  2. The gateway calling the backend.

These are not necessarily the same authentication mechanism.

Client-to-gateway authentication

The client may authenticate to API Management using:

  • An API Management subscription key.
  • Microsoft Entra authentication.
  • OAuth.
  • Another supported API authentication method.

For example, a client may send an API Management subscription key in a header expected by the API definition or policy.

Gateway-to-backend authentication

The gateway can authenticate to AI backends using:

  • Managed identity.
  • API keys.
  • Provider-specific credentials.
  • Other supported authentication mechanisms.

Managed identity is often preferable because it avoids embedding long-lived API keys in applications or configuration files. The managed identity must have the appropriate role or permissions on the backend.

Managed identity benefits

Using managed identity can:

  • Avoid storing API keys in application code.
  • Reduce credential rotation requirements.
  • Integrate with Microsoft Entra access control.
  • Support centralized identity governance.
  • Reduce the risk of accidental credential exposure.

However, managed identity does not automatically grant access. The identity must still be authorized on the target resource.


API Management Policies

API Management policies are XML-based rules that execute in the gateway.

Policies can operate on:

  • Inbound requests.
  • Backend requests.
  • Backend responses.
  • Outbound responses.
  • Errors.

They can validate, transform, secure, route, or limit API traffic. API Management policies are different from Azure Policy. API Management policies run at API request time, while Azure Policy evaluates and governs Azure resources.

Important AI Gateway policies

Rate limiting

Rate limiting restricts the number of requests a client can make during a defined period.

Example uses include:

  • Limiting requests per application.
  • Limiting requests per user.
  • Preventing excessive tool calls.
  • Protecting backend capacity.
  • Reducing abuse.

A rate-limit-by-key policy can use a key such as:

  • Client IP address.
  • Subscription key.
  • Application identifier.
  • User identifier.
  • Custom request value.

The key should be selected carefully. IP-based limiting may be inappropriate when many users share the same outbound address.

Token quotas

AI model consumption is often measured in tokens rather than only requests.

Token quotas can help control:

  • Cost.
  • Capacity.
  • Fair usage.
  • Project-level consumption.
  • Large prompt abuse.
  • Excessive response generation.

A request-per-minute limit alone may not prevent a client from sending extremely large prompts. Token-based controls are therefore important for AI workloads.

IP filtering

IP filtering can restrict requests to trusted networks or addresses.

For example, an organization may allow gateway access only from:

  • Corporate networks.
  • Private application subnets.
  • Approved build environments.
  • Trusted automation services.

IP filtering should not be treated as a replacement for identity-based authentication. Network location alone is not sufficient to establish who is authorized to use an AI service.

Authentication and authorization

Policies can validate tokens, inspect claims, and enforce access rules.

For example, a policy may:

  • Validate a Microsoft Entra token.
  • Check the token audience.
  • Restrict access to specific application IDs.
  • Require a subscription key.
  • Reject unauthenticated requests.
  • Route different consumers to different backends.

Content safety

API Management can apply policies that integrate with Azure AI Content Safety to moderate prompts or responses.

Content safety controls may help detect or block content such as:

  • Hate.
  • Violence.
  • Sexual content.
  • Self-harm content.
  • Other unsafe material, depending on the configured policy and service capabilities.

Content safety is not a complete AI security solution. It should be combined with identity, authorization, data protection, logging, and runtime controls.

Correlation IDs

A correlation ID allows related requests to be traced across systems.

A gateway can add a unique identifier to a request so that administrators can correlate:

  • Client requests.
  • Gateway logs.
  • Backend requests.
  • Application logs.
  • Security investigations.

Correlation IDs are particularly useful when an application invokes multiple models or tools during one user interaction.

Request and response transformation

Policies can modify requests or responses, including:

  • Headers.
  • URLs.
  • Query parameters.
  • Payloads.
  • Backend routing.
  • Response formatting.

Transformations should be used carefully with AI APIs because changing required headers or payload structures can cause model or tool calls to fail.


Governing MCP Tools

The Model Context Protocol allows agents to interact with external tools and data sources.

Examples of MCP tools include:

  • Search tools.
  • File access tools.
  • Database tools.
  • Business application tools.
  • Automation tools.
  • Custom enterprise tools.

Routing MCP traffic through an AI Gateway provides a centralized point for:

  • Authentication.
  • Rate limiting.
  • IP restrictions.
  • Audit logging.
  • Routing.
  • Policy enforcement.

The gateway can apply controls without requiring changes to the MCP server or agent code.

Important MCP limitations

For the Foundry-integrated MCP gateway experience:

  • The feature is in preview.
  • Only eligible MCP tools are routed through the gateway.
  • Existing tools may not be automatically migrated.
  • Tools using managed OAuth may not be eligible for the same routing flow.
  • API Management policies are configured in Azure API Management.
  • Gateway logs may not contain complete tool-level traces.
  • MCP server logs may still be required for detailed tool investigation.

If a tool was created before the gateway was connected, it may continue to call the MCP server directly. In that situation, recreate the tool after the gateway is connected if the tool is eligible for gateway routing.


Monitoring and Observability

An AI Gateway provides a central location for monitoring AI traffic.

Useful telemetry may include:

  • Request counts.
  • Response codes.
  • Latency.
  • Backend failures.
  • Token usage.
  • Rate-limit events.
  • Authentication failures.
  • Policy violations.
  • Correlation IDs.
  • Model or backend usage.
  • Gateway errors.

Telemetry can be viewed through API Management monitoring capabilities and, where configured, Microsoft Foundry or Application Insights.

Monitoring helps answer questions such as:

  • Which applications are using a model?
  • Which projects consume the most tokens?
  • Are requests being rejected?
  • Is a backend unavailable?
  • Are clients exceeding quotas?
  • Are unusual traffic patterns occurring?
  • Are policies blocking legitimate workloads?
  • Are sensitive tools being called unexpectedly?

Gateway telemetry should be combined with application, model, agent, and backend logs because the gateway may not capture every detail of an AI interaction.


Security Design Considerations

Use a single governed entry point

Where practical, route approved AI traffic through the gateway rather than allowing every application to call model endpoints directly.

This improves consistency and visibility.

Avoid exposing backend credentials

Prefer managed identity or centrally managed credentials rather than embedding API keys in application code.

Apply least privilege

The gateway’s identity should have only the permissions required to access the configured backend.

The client should also receive only the permissions required to call the gateway APIs.

Separate environments

Use separate configurations or gateways for:

  • Development.
  • Testing.
  • Staging.
  • Production.

This helps prevent development applications from accessing production models or sensitive tools.

Apply quotas by project or consumer

Quotas should reflect business requirements. A shared global quota may allow one application to consume capacity needed by other teams.

Restrict sensitive tools

Tools that can:

  • Modify databases.
  • Send email.
  • Change permissions.
  • Deploy resources.
  • Access confidential data.
  • Execute commands.

should receive stronger controls than read-only tools.

Protect private backends

If a Foundry resource is private, ensure the gateway has appropriate private connectivity. Do not assume that associating the resources automatically solves networking requirements.

Test policies before enforcement

Use testing and staged rollout to ensure that policies do not:

  • Remove required authentication headers.
  • Break model payloads.
  • Block legitimate clients.
  • Prevent required tool calls.
  • Cause unexpected latency.
  • Interfere with streaming responses.

Example Scenario

A company has three AI applications:

  • A customer-service assistant.
  • A financial reporting application.
  • An internal research assistant.

Each application currently calls a model endpoint directly.

The security team deploys Azure API Management as an AI Gateway and configures:

  1. Microsoft Entra authentication for approved applications.
  2. Managed identity authentication from the gateway to the model backend.
  3. Separate API products or subscriptions for each application.
  4. Token quotas for each project.
  5. Rate limits to prevent excessive requests.
  6. IP restrictions for internal applications.
  7. Content safety checks.
  8. Correlation IDs for tracing.
  9. Monitoring and alerts for failures and abnormal usage.

The financial reporting application receives a lower quota but stronger access restrictions because it processes sensitive information. The research assistant is allowed to use several approved models but cannot access production business tools.

This design provides centralized security while allowing different applications to have different access and usage policies.


Common Exam Comparisons

CapabilityPrimary purpose
Azure API Management AI GatewayGovern and secure traffic to AI models, agents, and tools
Microsoft FoundryDevelop, deploy, and operate AI resources
Microsoft Entra IDAuthenticate identities and authorize access
Managed identityProvide Azure-managed authentication without embedded credentials
API Management policyApply runtime request and response controls
Azure PolicyGovern Azure resource configuration and compliance
Azure AI Content SafetyDetect or moderate unsafe content
Application InsightsMonitor application and gateway-related telemetry
Microsoft SentinelCentralize security events and investigation workflows
Agent identityRepresent an AI agent as an identity
MCP serverExpose tools or context to compatible AI clients

Exam Tips

Remember these key points:

  • An AI Gateway is implemented using Azure API Management capabilities.
  • The gateway is a control point between AI consumers and AI backends.
  • It can govern models, agents, and eligible MCP tools.
  • The gateway is not a replacement for Microsoft Entra authentication.
  • Client-to-gateway authentication and gateway-to-backend authentication are separate concerns.
  • Managed identity can reduce the need to store backend API keys.
  • API Management policies are runtime rules, not Azure Policy definitions.
  • Rate limits control request volume; token quotas control AI consumption.
  • Content safety policies address unsafe content but do not replace identity security.
  • Existing MCP tools may not automatically begin using a newly connected gateway.
  • API Management policies are configured in API Management, not necessarily in the Foundry portal.
  • Private Foundry resources require compatible private connectivity from the gateway.
  • Preview features may have changing capabilities, limits, and supported regions.
  • Monitoring should include gateway telemetry and backend or application logs.
  • Always verify that traffic actually passes through the gateway after configuration.

Practice Exam Questions

Question 1

What is the primary purpose of using Azure API Management as an AI Gateway for Microsoft Foundry?

A. To replace Microsoft Entra ID
B. To train foundation models
C. To provide a centralized point for securing, governing, and monitoring AI traffic
D. To permanently store model training data

Answer: C

Explanation: An AI Gateway provides a centralized control point between AI consumers and backends. It can enforce authentication, quotas, rate limits, routing, monitoring, and other policies. It does not replace Microsoft Entra ID or train models.


Question 2

An organization wants the AI Gateway to authenticate to a Microsoft Foundry backend without storing an API key in application code. Which option should it consider?

A. Managed identity
B. Azure resource lock
C. Public IP address filtering only
D. A storage account access key

Answer: A

Explanation: Managed identity allows the gateway to authenticate to supported Azure services without embedding long-lived credentials in application code. The identity must still be granted the required permissions on the backend.


Question 3

Which statement correctly describes Azure API Management policies?

A. They are Azure resource-compliance definitions evaluated by Azure Policy
B. They are runtime rules that can validate, transform, secure, limit, or route API requests and responses
C. They are Microsoft Entra role assignments
D. They are model-training instructions

Answer: B

Explanation: API Management policies execute in the gateway and can control inbound requests, backend requests, responses, and errors. They are different from Azure Policy, which governs Azure resource configuration.


Question 4

A company wants to prevent one Foundry project from consuming all available model capacity. Which control is most directly relevant?

A. Azure Bastion
B. Microsoft Entra access reviews
C. Token quotas
D. Azure resource locks

Answer: C

Explanation: Token quotas limit AI consumption based on token usage. They are useful for cost control, capacity management, and preventing one project from monopolizing model capacity.


Question 5

An administrator connects an AI Gateway to a Foundry resource. An MCP tool created several weeks earlier continues to call the MCP server directly. What is the most likely explanation?

A. API Management cannot govern MCP tools
B. The tool must be recreated after the gateway is connected if it is eligible for gateway routing
C. The Foundry resource must be deleted
D. The MCP server must be converted into a virtual machine

Answer: B

Explanation: In the preview Foundry MCP gateway integration, gateway routing is applied when eligible tools are created after the gateway is connected. Existing tools are not automatically migrated.


Question 6

Which control is most appropriate for restricting AI Gateway requests to approved corporate networks?

A. IP filtering
B. Token quota
C. Model fine-tuning
D. Agent blueprint inheritance

Answer: A

Explanation: IP filtering can allow or deny requests based on source IP addresses or ranges. It should supplement, not replace, identity-based authentication and authorization.


Question 7

An organization wants to trace a request from an application through API Management to the model backend. Which feature is most useful?

A. Azure resource locks
B. Correlation IDs
C. Disk encryption
D. Azure Policy remediation

Answer: B

Explanation: Correlation IDs provide a common identifier that can be recorded in gateway, application, and backend logs, making it easier to trace a request across multiple components.


Question 8

A Foundry resource has public network access disabled. What must be verified before associating it with an AI Gateway?

A. The gateway has an appropriate private network path to the Foundry resource
B. The model has been fine-tuned
C. All clients use anonymous access
D. The gateway is deployed outside Azure

Answer: A

Explanation: A private Foundry resource requires compatible private connectivity from the API Management instance. The gateway must be able to reach the backend through the required private networking configuration.


Question 9

Which statement best describes the difference between rate limiting and token quotas?

A. Rate limiting controls request volume, while token quotas control AI token consumption
B. Rate limiting authenticates users, while token quotas encrypt traffic
C. Rate limiting governs Azure resources, while token quotas manage virtual machines
D. Rate limiting disables models, while token quotas create agents

Answer: A

Explanation: Rate limiting restricts the number of requests over a period. Token quotas restrict the amount of model consumption, which is important because requests can vary greatly in prompt and response size.


Question 10

An administrator applies an API Management policy that removes a required authentication header before forwarding the request to an MCP server. What is the likely result?

A. The MCP server automatically repairs the header
B. The request is converted into a model-training job
C. The request may fail because required authentication information was removed
D. The gateway automatically grants anonymous access

Answer: C

Explanation: API Management policies can modify headers, but removing a required authentication header can cause the backend or MCP server to reject the request. Policies must be tested carefully to avoid breaking required authentication and protocol behavior.


Final Exam Point

A very important exam theme is that Azure API Management provides the enforcement and governance layer, while Microsoft Foundry provides the AI resource and application environment. Together, they allow organizations to centralize access control, usage management, monitoring, and security for AI workloads.


Go to the SC-500 Exam Prep Hub main page