Tag: Azure API Management

Implement security policies for back-end API protection by using API Management (SC-500 Exam Prep)

This post is a part of the "SC-500: Implementing End-to-End Security Controls for Cloud and AI Workloads" Exam Prep Hub.
This topic falls under these sections:
Secure compute (20–25%)
   --> Implement security for application platform services
      --> Implement security policies for back-end API protection by using API Management


Note that there are 10 practice questions (with answers) at the end of each section to help you solidify your knowledge of the material. Also, there are 4 practice tests with 30 questions each available from the hub's main page below the exam topics section.

Introduction

APIs are a critical part of modern cloud applications. They allow web applications, mobile applications, business systems, AI applications, and other services to communicate with back-end resources.

However, exposing an API creates a significant security responsibility. Organizations must control:

  • Who can call an API
  • How callers authenticate
  • What callers are authorized to do
  • How frequently APIs can be called
  • Where requests originate
  • What traffic is allowed to reach back-end services
  • How API credentials are protected
  • How connections between API Management and back-end services are secured
  • How potentially malicious or malformed requests are handled

Azure API Management (APIM) provides an API gateway that can enforce security policies between API consumers and back-end APIs.

The SC-500 exam specifically expects knowledge of securing back-end APIs with API Management, including subscription keys, JWT validation, OAuth 2.0, IP filtering, rate limiting, virtual network integration, and mutual TLS (mTLS).

A useful way to think about API Management security is:

                    API Consumer
                         |
                         v
              +---------------------+
              | Azure API Management|
              |       Gateway       |
              +---------------------+
                 |       |       |
          Authentication  |   Network
          Authorization   |   Controls
                 |        |
                 v        v
             Policies   Filtering
                 |
                 v
              Backend API

API Management policies execute in the gateway and can inspect, validate, transform, route, and control API requests and responses. Microsoft currently provides more than 75 built-in policies covering scenarios such as authentication, rate limiting, filtering, transformation, and validation.


1. Understanding the API Management Security Model

API Management sits between the API consumer and the back-end API.

Without API Management:

Client
|
v
Backend API

With API Management:

Client
|
v
Azure API Management
|
|-- Authenticate
|-- Authorize
|-- Validate
|-- Rate limit
|-- Filter IPs
|-- Transform
|-- Log/monitor
|
v
Backend API

This creates a centralized enforcement point.

Instead of implementing every security control independently inside every API, organizations can implement common controls at the API gateway.

Important distinction

API Management policies are runtime API policies.

They should not be confused with Azure Policy.

FeaturePurpose
API Management policyControls API requests/responses at runtime
Azure PolicyGoverns Azure resources and enforces organizational compliance
Azure RBACControls who can manage Azure resources
Microsoft Entra IDProvides identity and authentication
Network Security GroupControls network traffic
Azure FirewallProvides centralized network traffic filtering

Microsoft specifically distinguishes API Management policies from Azure Policy: APIM policies execute in the gateway, while Azure Policy evaluates and enforces compliance for Azure resources.


2. API Management Policy Scopes

Policies can be applied at different levels.

Common scopes include:

  • Global/API Management instance
  • Product
  • API
  • Operation

This allows security controls to be applied broadly or narrowly.

For example:

API Management
|
+-- Global policy
|
+-- Product policy
|
+-- API policy
|
+-- Operation policy

Why scope matters

Suppose every API requires a particular security header.

A global policy may make sense.

Suppose only the /payments API requires a particular JWT claim.

An API-level policy may be more appropriate.

Suppose only POST /payments requires an additional validation rule.

An operation-level policy may be appropriate.

Best practice

Apply security controls at the narrowest practical scope when the requirement is specific to a particular API or operation, while using broader scopes for controls that should consistently apply across the environment.


3. Subscription Keys

Azure API Management supports subscription keys as a mechanism for controlling access to APIs.

A subscription represents a relationship between an API consumer and one or more APIs or products.

The consumer can provide a subscription key with an API request.

Conceptually:

Client
|
| Subscription key
v
API Management
|
| Validate subscription
v
Backend API

Subscription keys provide a way to identify and control API consumers.

They can also be useful for:

  • Tracking API consumption
  • Managing subscriptions
  • Applying quotas or rate limits
  • Controlling access to products/APIs

Important exam distinction

A subscription key should not automatically be considered equivalent to strong user authentication.

A subscription key primarily identifies an API subscription.

It does not inherently prove the identity of an individual user in the same way that a Microsoft Entra access token can.

For applications requiring user or workload identity, OAuth 2.0 and JWT validation may be more appropriate.


4. Subscription Keys and Products

API Management products provide an important organizational mechanism for grouping APIs and managing consumer access.

For example:

Product: Customer APIs
|
+-- Customer API
+-- Orders API
+-- Account API

A subscription can provide access to a product.

This makes it possible to manage access to groups of APIs rather than independently configuring every API consumer for every individual API.

Example

A company might create:

Internal APIs

  • Employee API
  • HR API
  • Finance API

Partner APIs

  • Orders API
  • Inventory API
  • Shipping API

Different consumers can receive different subscriptions and access different products.


5. JWT-Based Authentication

JSON Web Tokens, or JWTs, are commonly used to authenticate API callers.

A JWT can contain claims about:

  • The caller
  • The issuer
  • The intended audience
  • Expiration
  • Permissions or scopes
  • Other identity information

A typical flow is:

Client
|
| Authorization: Bearer <JWT>
v
API Management
|
| Validate JWT
|
+---- Invalid ----> Reject
|
+---- Valid ------> Backend API

API Management provides a validate-jwt policy that can validate JWT tokens. The policy can verify aspects such as token issuer, audience, signature, expiration, and claims depending on its configuration.


6. Validate JWT Policy

The validate-jwt policy is one of the most important policies for SC-500.

A simplified policy looks conceptually like:

<validate-jwt header-name="Authorization"
require-scheme="Bearer">
...
</validate-jwt>

The actual configuration must specify the appropriate issuer/signing keys/audience and other requirements for the identity provider and API.

The policy can validate that:

  • A token exists
  • The token is properly signed
  • The token was issued by an expected issuer
  • The token is intended for the expected audience
  • The token has not expired
  • Required claims are present

Why audience matters

Consider a token issued for:

Audience = API-A

An attacker attempts to use that token against:

API-B

If API-B validates the audience correctly, it can reject the token.

This is an important security control.


7. Microsoft Entra ID and JWT Validation

Microsoft Entra ID is frequently used as the identity provider for API authentication.

The flow can look like:

User/Application
|
| Authenticate
v
Microsoft Entra ID
|
| Access token
v
API Management
|
| Validate JWT
v
Backend API

API Management can validate Microsoft Entra-issued tokens using JWT validation policies. Current API Management policy support includes a dedicated validate Microsoft Entra token policy as well as the general JWT validation policy.

Exam concept

If the question asks:

“How can API Management verify that an incoming request contains a valid Microsoft Entra access token?”

Think:

JWT/token validation policy.


8. Claims-Based Authorization

Authentication answers:

Who are you?

Authorization answers:

What are you allowed to do?

JWT claims can help API Management enforce authorization requirements.

For example, an API may require:

scope = Orders.Read

A more privileged API may require:

scope = Orders.Write

API Management can inspect token claims and reject requests that do not satisfy the required conditions.

This allows security decisions to occur before the request reaches the backend.


9. OAuth 2.0

OAuth 2.0 provides a standardized framework for delegated authorization.

A simplified flow is:

Client
|
| Request authorization
v
Identity Provider
|
| Access token
v
Client
|
| Bearer token
v
API Management
|
| Validate
v
Backend API

API Management can be configured to work with OAuth 2.0 authorization servers, including Microsoft Entra ID. Microsoft currently recommends the Microsoft identity platform v2 endpoint for OAuth scenarios where applicable.

OAuth vs subscription key

CapabilitySubscription keyOAuth/JWT
API subscription identificationYesNot its primary purpose
User identityLimitedStronger identity model
Token expirationNo inherent token conceptYes
Claims/scopesNoYes
Delegated authorizationNoYes
Microsoft Entra integrationPossible as part of broader designStrong integration

10. OAuth 2.0 for Backend APIs

API Management can also help manage credentials and authorization for calls from API Management to backend services.

This is an important distinction.

There are potentially two authentication relationships:

Consumer
|
| Authenticate
v
API Management
|
| Authenticate
v
Backend API

The authentication method used by the consumer does not necessarily have to be the same authentication method used by API Management when it calls the backend.

API Management’s current credential manager can manage OAuth 2.0 connections to backend APIs and can acquire, cache, and refresh tokens for supported scenarios.


11. Managed Identity for Backend Authentication

API Management can authenticate to supported backend services using a managed identity.

This can reduce the need to store credentials such as:

  • Client secrets
  • Passwords
  • Long-lived access tokens

Conceptually:

API Management
|
| Managed identity
v
Azure Resource / Backend

The API Management policy reference currently includes an authenticate with managed identity policy for authenticating to backend services.

Why this is important

Managed identity follows the principle:

Don’t store a secret when Azure can provide an identity.

It is therefore generally preferable to hard-coded credentials.


12. IP Filtering

API Management can restrict API access based on the caller’s IP address.

For example:

Corporate network
10.20.0.0/16
|
v
API Management
|
+---- Allow

And:

Unknown source
203.0.113.0/24
|
v
API Management
|
+---- Deny

API Management provides a policy to restrict callers by IP address or IP address range.

When IP filtering is useful

It can be appropriate for:

  • Internal APIs
  • Partner APIs
  • Administrative APIs
  • Known integration platforms
  • Restricted enterprise networks

Important limitation

IP filtering should not be treated as a replacement for authentication.

An IP address identifies a network source, not necessarily an individual user.


13. Rate Limiting

An API can become unavailable or expensive when clients send excessive requests.

API Management can apply rate limiting policies.

For example:

Client
|
| 100 requests
|-------------------->
|
API Management
|
| Threshold = 50
|
+---- Additional requests rejected

Rate limiting can protect:

  • Application capacity
  • Backend systems
  • APIs with expensive operations
  • Shared services
  • APIs susceptible to abuse

API Management policies include rate-limiting capabilities for controlling incoming calls.


14. Rate Limiting vs Quotas

Rate limiting and quotas are related but different concepts.

Rate limiting

Controls how many requests can occur during a relatively short period.

Example:

No more than 100 calls per minute.

Quota

Controls consumption over a longer period.

Example:

No more than 100,000 calls per month.

Conceptually:

Rate limit:
100 calls / minute
Quota:
100,000 calls / month

Exam tip

If the scenario describes burst/excessive requests over a short period, think rate limiting.

If it describes total consumption over a longer period, think quota.


15. IP Filtering + Rate Limiting

Security policies can be combined.

For example:

Incoming request
|
v
IP filter
|
+---- Unauthorized source --> Reject
|
v
JWT validation
|
+---- Invalid token -------> Reject
|
v
Rate limit
|
+---- Too many requests ---> Reject
|
v
Backend API

This is an example of defense in depth.

No single policy has to solve every security problem.


16. Request and Response Validation

API Management can validate request and response content.

The validate-content policy can validate the size and/or content of request or response bodies against supported schemas, including JSON and XML.

This can help prevent unexpected or malformed content from reaching an API.

For example:

Client
|
| JSON request
v
API Management
|
| Validate schema
|
+---- Invalid --> Reject
|
v
Backend API

This is particularly useful for APIs with strict data contracts.


17. Mutual TLS (mTLS)

Mutual TLS, or mTLS, provides certificate-based authentication between communicating parties.

With normal TLS:

Client <---- TLS ----> Server

The server proves its identity to the client.

With mTLS:

Client <---- Mutual TLS ----> Server
| |
+-- Client certificate |
|
Server certificate

Both sides authenticate using certificates.

API Management supports policies for validating client certificates and authenticating to backend services using client certificates.


18. mTLS for Backend APIs

One particularly important use case is securing the connection between API Management and a backend API.

For example:

API Consumer
|
| HTTPS
v
API Management
|
| mTLS
v
Secure Backend API

The backend can require API Management to present a client certificate.

API Management supports an authentication-certificate policy for authenticating with a backend service using a client certificate.

When to use mTLS

mTLS is particularly appropriate when:

  • A backend requires certificate-based client authentication.
  • Strong machine-to-machine authentication is required.
  • An organization operates partner APIs requiring certificates.
  • Traditional API keys are insufficient.
  • The environment has established PKI infrastructure.

19. Network Security for API Management

API-level policies are only one layer of security.

API Management can also be integrated with Azure networking capabilities.

Current API Management networking options include:

  • Virtual network injection
  • Virtual network integration
  • Inbound private endpoints
  • Private Link
  • Network security controls

The exact capabilities depend on the API Management tier.


20. Virtual Network Integration

Virtual network integration can allow API Management to make outbound requests to APIs hosted in a connected virtual network.

For example:

                    Azure VNet
             +----------------------+
             |                      |
API Management -----> Private API   |
             |                      |
             +----------------------+

For current Standard v2 and Premium v2 configurations, virtual network integration supports outbound connectivity to APIs isolated in a connected or peered virtual network. However, this configuration does not by itself make the API Management gateway private; the gateway and developer portal remain publicly accessible unless additional controls are configured.

Exam trap

Do not assume:

VNet integration = private inbound API Management endpoint.

VNet integration in these v2 scenarios primarily provides outbound connectivity.


21. Inbound Private Endpoint

An inbound private endpoint provides private connectivity to the API Management instance.

Conceptually:

Private Client
|
v
Private Endpoint
|
v
Azure API Management
|
v
Backend API

The private endpoint receives an IP address from the virtual network.

Traffic can then travel through Azure Private Link rather than being exposed through the public internet.

Microsoft’s current documentation states that the inbound private endpoint is for inbound traffic and that public network access can be disabled after configuring a private endpoint.

Important distinction

CapabilityPrimary purpose
VNet integrationOutbound connectivity
Inbound private endpointPrivate inbound access
VNet injectionNetwork isolation of APIM, depending on tier/configuration

This distinction is highly relevant for exam questions.


22. Virtual Network Injection

Some API Management tiers support virtual network injection.

In this architecture, the API Management instance can be deployed into a delegated subnet.

For supported configurations, this can provide stronger network isolation for the API Management gateway and its backend connectivity.

The current Premium v2 documentation describes virtual network injection as providing both inbound and outbound network connectivity through the virtual network and is intended for scenarios where both the API Management instance and backend APIs need isolation.

Exam concept

When a scenario emphasizes:

“Isolate the API Management instance itself inside a private network.”

Think about virtual network injection/private networking, depending on the APIM tier and architecture.


23. Securing the Backend

Securing API Management does not automatically secure the backend API.

Consider:

Internet
|
v
API Management
|
v
Backend API

The backend should also be protected.

Potential controls include:

  • Private networking
  • Authentication
  • Managed identity
  • mTLS
  • IP restrictions
  • Network security groups where applicable
  • Firewall controls
  • Private endpoints
  • Network security perimeter capabilities where appropriate

The objective should be to ensure that the backend accepts traffic from legitimate API Management paths rather than becoming an independent publicly exposed API.


24. API Management as a Security Gateway

A mature architecture often uses APIM as the central policy enforcement point.

For example:

                       Clients
                          |
             +------------+------------+
             |            |            |
          Web App      Mobile App    Partner
             |            |            |
             +------------+------------+
                          |
                          v
               +---------------------+
               | Azure API Management|
               |                     |
               | Subscription keys   |
               | JWT validation      |
               | OAuth 2.0           |
               | IP filtering        |
               | Rate limiting       |
               | Content validation  |
               | mTLS                |
               +---------------------+
                          |
                          v
                 Private Backend APIs

This architecture centralizes common security requirements.


25. AI APIs and API Management

API Management can also be used as a security and governance layer for AI model endpoints.

This is increasingly important because AI applications often expose APIs to:

  • Language models
  • AI agents
  • AI applications
  • Retrieval systems
  • External consumers
  • Internal applications

The current SC-500 learning module explicitly includes AI Gateway capabilities for securing and governing AI model endpoints.

This extends the API gateway concept beyond traditional REST APIs.

An AI gateway can help organizations apply centralized policies to AI traffic, such as:

  • Authentication
  • Authorization
  • Rate limiting
  • Content controls
  • Usage governance
  • Monitoring

26. AI Gateway and Content Safety

API Management’s current policy framework includes AI-specific capabilities, including enforcing content-safety checks on LLM requests by sending prompts to Azure AI Content Safety before they are sent to the backend LLM.

Conceptually:

AI Application
|
v
API Management AI Gateway
|
+---- Authentication
|
+---- Rate limiting
|
+---- Content safety
|
+---- Governance
|
v
LLM / AI Backend

This is particularly relevant to the evolving SC-500 emphasis on securing cloud and AI workloads.


27. A Layered API Security Model

A useful way to remember the complete approach is to divide security into layers.

LayerExample control
IdentityMicrosoft Entra ID
AuthenticationOAuth 2.0 / JWT
SubscriptionSubscription keys
AuthorizationJWT claims/scopes
NetworkVNet/private endpoint
Source filteringIP filtering
Abuse protectionRate limiting
ContentSchema/content validation
Machine authenticationmTLS
Backend authenticationManaged identity
AI securityAI Gateway/content safety
MonitoringAPI Management/Azure Monitor

No single control is sufficient for every API.


28. Example: Secure a Partner API

Consider a company exposing an Orders API to a business partner.

Requirements:

  • Only the partner can access the API.
  • Requests must contain valid Microsoft Entra tokens.
  • The partner’s network has known IP ranges.
  • Excessive API requests must be limited.
  • The backend must not be publicly exposed.
  • API Management must authenticate to the backend.

A possible design is:

Partner
|
| Microsoft Entra token
v
API Management
|
+-- Validate JWT
|
+-- IP filtering
|
+-- Rate limiting
|
+-- Subscription management
|
v
Private Backend
^
|
Managed Identity

This provides multiple independent controls.

If an attacker obtains an IP address within an allowed range but lacks a valid token, JWT validation can reject the request.

If a legitimate partner begins sending excessive traffic, rate limiting can protect the backend.

If the backend is private, network controls can prevent direct public access.


29. Example: Secure a Machine-to-Machine API

Suppose an enterprise application calls a sensitive internal API.

Requirements:

  • No interactive user sign-in
  • Strong workload identity
  • Private network connectivity
  • Certificate-based backend authentication

A suitable architecture could use:

Enterprise Application
|
| Token / approved authentication
v
API Management
|
| Private network
|
| mTLS
v
Internal Backend

The exact combination depends on the identity architecture, but the key concept is that machine identity, network isolation, and certificate authentication solve different security problems.


30. Common API Management Security Mistakes

Mistake 1: Treating a subscription key as complete authentication

A subscription key identifies API subscription access but should not automatically be treated as equivalent to strong user or workload identity.

Better: Use OAuth 2.0/JWT validation when identity and claims-based authorization are required.


Mistake 2: Using IP filtering instead of authentication

An IP address does not establish who the caller is.

Better: Combine IP filtering with strong authentication where appropriate.


Mistake 3: Confusing VNet integration with a private inbound gateway

In current v2 configurations, VNet integration primarily enables outbound connectivity to private backends while the gateway remains publicly accessible.

Better: Use the appropriate private endpoint or VNet injection architecture when private inbound access is required.


Mistake 4: Exposing the backend publicly

Putting API Management in front of an API does not automatically make the backend private.

Better: Restrict backend access and use private networking where appropriate.


Mistake 5: Hard-coding backend secrets

Storing long-lived credentials in application configuration increases security risk.

Better: Prefer managed identity or appropriate secure credential-management mechanisms.


Mistake 6: Applying rate limiting only inside application code

Application-level rate limiting may require every API to independently implement and maintain the same logic.

Better: Use API Management policies for centralized API traffic controls.


Mistake 7: Using broad policies when narrow policies are sufficient

A global security rule can unintentionally affect APIs that do not require it.

Better: Choose the appropriate policy scope.


31. API Management Security Decision Matrix

RequirementAPIM capability
Identify API subscriptionSubscription key
Authenticate users/workloadsOAuth 2.0 / JWT
Validate Microsoft Entra access tokenJWT/Entra token validation
Check token claims/scopesJWT validation/authorization logic
Block specific source IPsIP filtering
Limit requests per periodRate limiting
Limit longer-term consumptionQuotas
Validate request bodyContent/schema validation
Authenticate backend with Azure identityManaged identity
Authenticate backend with certificatemTLS/client certificate
Private inbound APIM accessPrivate endpoint
Reach private backendVNet integration/injection as supported
Isolate APIM and backendAppropriate VNet/private architecture
Protect AI endpointsAI Gateway policies
Apply LLM content safety checksAI/content-safety policy
Centralize API securityAPI Management policies

32. SC-500 Exam-Focused Comparison

Several capabilities can look similar in scenario questions.

If the question says…Think…
“Identify the API consumer’s subscription”Subscription key
“Validate an access token”JWT validation
“Microsoft Entra authentication”OAuth 2.0 / JWT
“Validate required token claims”JWT validation
“Block calls from a specific IP range”IP filtering
“Prevent excessive requests”Rate limiting
“Limit total API usage over a period”Quota
“Validate JSON/XML body”Content validation
“Authenticate APIM to backend without a secret”Managed identity
“Certificate-based backend authentication”mTLS/client certificate
“Allow private clients to reach APIM”Private endpoint
“Allow APIM to reach a private backend”VNet integration/injection
“Keep APIM and backend isolated”Private networking/VNet architecture
“Secure AI model endpoints”AI Gateway

33. Key Takeaways

For SC-500, remember these principles:

  1. API Management is a centralized API gateway and policy enforcement point.
  2. Subscription keys provide subscription-based API access and identification.
  3. JWT validation verifies access tokens and their claims.
  4. OAuth 2.0 provides a standardized authorization framework.
  5. Microsoft Entra ID can be used as the identity provider for OAuth/JWT scenarios.
  6. IP filtering controls where API requests originate.
  7. Rate limiting controls excessive short-term request activity.
  8. Quotas control longer-term API consumption.
  9. Managed identity can authenticate API Management to supported backend resources without storing credentials.
  10. mTLS provides certificate-based authentication between communicating parties.
  11. Content validation can reject requests or responses that do not conform to expected schemas or size requirements.
  12. VNet integration and private endpoints solve different networking problems.
  13. An inbound private endpoint provides private access to the API Management gateway.
  14. VNet integration in current v2 tiers primarily enables outbound access to private backends; it does not automatically make the gateway private.
  15. AI Gateway extends API Management security and governance to AI model endpoints.
  16. Security controls should be layered rather than relying on one mechanism.

Practice Exam Questions

Question 1

A company exposes several APIs through Azure API Management. The company wants to identify API consumers and control which APIs they can access through an API subscription.

Which capability should the security engineer configure?

A. Azure Private Endpoint

B. Subscription keys

C. Client certificates

D. IP filtering

Answer: B

Explanation

Subscription keys are designed to provide subscription-based access to APIs and identify API consumers within the API Management subscription model.

Private endpoints address network connectivity, client certificates provide certificate-based authentication, and IP filtering restricts traffic based on source IP addresses.


Question 2

An organization exposes a sensitive API through API Management. The API should accept requests only when the caller presents a valid Microsoft Entra access token intended for the API and containing the required claims.

Which capability should be implemented?

A. Rate limiting

B. Subscription key validation only

C. IP filtering

D. JWT/token validation

Answer: D

Explanation

A JWT validation policy can validate the access token’s signature, issuer, audience, expiration, and required claims according to the configured policy.

A subscription key alone does not provide equivalent identity and claims-based authorization.


Question 3

An API Management instance must call an Azure backend service. The security team does not want to store a client secret or password in the API configuration.

Which authentication mechanism is most appropriate?

A. Managed identity

B. IP filtering

C. Subscription key

D. Anonymous access

Answer: A

Explanation

A managed identity allows API Management to authenticate to supported Azure resources without storing a long-lived credential in the application or API policy.

This supports a secretless authentication model and aligns with least-credential principles.


Question 4

A public API is being abused by a client that sends thousands of requests within a short period. The organization wants API Management to automatically restrict the number of requests the client can make during a defined interval.

Which capability should be used?

A. OAuth 2.0

B. mTLS

C. Rate limiting

D. Private endpoint

Answer: C

Explanation

Rate limiting is designed to control the number of calls received during a defined period.

OAuth 2.0 addresses authorization, mTLS provides certificate-based authentication, and private endpoints address network connectivity.


Question 5

An organization has API Management integrated with a virtual network. The API Management instance must call an API hosted on a private subnet. However, the security team also wants API consumers to continue reaching the API Management gateway through its public endpoint.

Which capability best satisfies the networking requirement?

A. Disable all public access to API Management

B. Configure outbound virtual network integration

C. Configure an API subscription key

D. Configure JWT validation

Answer: B

Explanation

Current Standard v2 and Premium v2 API Management configurations support virtual network integration for outbound requests to APIs isolated within a connected or peered virtual network.

Importantly, this does not by itself make the API Management gateway private; its gateway and developer portal can remain publicly accessible.


Question 6

A company wants only clients from a set of known corporate IP ranges to call a particular API through API Management.

Which policy should the security engineer use?

A. IP filtering

B. JWT validation

C. Content validation

D. Managed identity

Answer: A

Explanation

The API Management IP filtering capability can allow or deny calls based on specific IP addresses or address ranges.

However, IP filtering should generally complement rather than replace strong authentication.


Question 7

A company requires API Management to authenticate to a highly sensitive partner backend using a client certificate. The partner will reject requests unless API Management presents the expected certificate.

Which capability should be configured?

A. Subscription key

B. Rate limiting

C. OAuth authorization code

D. Client certificate authentication/mTLS

Answer: D

Explanation

mTLS/client certificate authentication is appropriate when a backend requires certificate-based authentication from API Management.

API Management provides an authentication-certificate policy for authenticating with a backend using a client certificate.


Question 8

A security engineer wants API Management to reject requests whose JSON body does not conform to the API’s expected schema.

Which capability should be used?

A. IP filtering

B. Subscription key validation

C. Content validation

D. Private endpoint

Answer: C

Explanation

The validate-content policy can validate the size and/or content of request or response bodies against supported schemas, including JSON and XML.

The other options address identity, network access, or subscription management.


Question 9

A company wants clients inside its private Azure network to connect to an API Management gateway without exposing the gateway through the public internet. The organization also wants to disable public network access to the API Management instance.

Which solution is most appropriate?

A. Inbound private endpoint

B. Rate limiting

C. Subscription key

D. JWT validation

Answer: A

Explanation

An inbound private endpoint provides private connectivity to the API Management instance using Azure Private Link and a private IP address from the virtual network.

Current API Management documentation states that public network access can be disabled after configuring an inbound private endpoint.


Question 10

A company is building an AI application that accesses an LLM through Azure API Management. The security team wants a centralized gateway that can apply authentication, traffic controls, and AI-specific protections before requests reach the model endpoint.

Which capability is most appropriate?

A. Azure Bastion

B. AI Gateway capabilities in API Management

C. Azure Network Security Group only

D. Azure Private DNS

Answer: B

Explanation

API Management’s AI Gateway capabilities extend gateway security and governance to AI model endpoints. Current policy capabilities include AI-specific controls such as content-safety checks for LLM requests, in addition to traditional API controls such as authentication and rate limiting.

Azure Bastion is designed for secure VM administration, NSGs provide network-level filtering, and Private DNS provides name resolution rather than centralized AI API governance.


Go to the SC-500 Exam Prep Hub main page

Configure and deploy AI Gateway in Azure API Management for Microsoft Foundry (SC-500 Exam Prep)

This post is a part of the "SC-500: Implementing End-to-End Security Controls for Cloud and AI Workloads" Exam Prep Hub.
This topic falls under these sections:
Secure compute (20–25%)
   --> Implement security for AI
      --> Configure and deploy AI Gateway in Azure API Management for Microsoft Foundry


Note that there are 10 practice questions (with answers) at the end of each section to help you solidify your knowledge of the material. Also, there are 4 practice tests with 30 questions each available from the hub's main page below the exam topics section.

Overview

Microsoft Foundry provides services for developing, deploying, and operating generative AI applications, models, and agents. As organizations adopt AI at scale, they need a controlled way to manage access to models and AI tools.

An AI Gateway uses Azure API Management to provide a governed entry point between applications or agents and AI backends. It can centralize authentication, authorization, traffic control, token usage, monitoring, routing, and security policies.

For SC-500, the important concept is that the AI Gateway is not simply another model endpoint. It is a security and governance layer placed between AI consumers and the services they access.

Exam focus: Understand how to connect Microsoft Foundry to Azure API Management, configure the gateway, import models or tools, apply policies, and verify that traffic is actually being mediated by the gateway.


What Is an AI Gateway?

An AI Gateway is a set of capabilities in Azure API Management that helps organizations manage AI-related backends.

These backends may include:

  • Models deployed in Microsoft Foundry.
  • Azure OpenAI deployments.
  • Other supported model providers.
  • OpenAI-compatible model endpoints.
  • Remote Model Context Protocol (MCP) servers.
  • Agent-to-agent APIs.
  • Custom AI services.
  • Self-hosted models and endpoints.

The AI Gateway extends the existing API Management gateway. It is not a completely separate gateway product. Existing API Management capabilities, including policies, authentication, routing, monitoring, and networking, are used to govern AI traffic.

Why use an AI Gateway?

Without a gateway, each application or agent may connect directly to an AI model or tool. This can lead to:

  • Duplicated authentication logic.
  • Inconsistent security policies.
  • Uncontrolled model consumption.
  • Difficulty enforcing quotas.
  • Limited visibility into usage.
  • Excessive exposure of backend endpoints.
  • Different teams implementing different controls.
  • Difficulty changing model providers.

An AI Gateway provides a centralized control point for these concerns.

For example, several applications might use different model deployments, but all requests can pass through API Management where the organization applies:

  • Authentication.
  • Authorization.
  • Rate limits.
  • Token quotas.
  • IP restrictions.
  • Content safety policies.
  • Request and response transformations.
  • Logging and metrics.
  • Backend routing.
  • Load balancing.
  • Caching, where appropriate.

AI Gateway Architecture

A typical architecture contains the following components:

  1. AI consumer
    • Application.
    • Copilot.
    • Agent.
    • Development tool.
    • Automated workload.
  2. Microsoft Foundry resource or project
    • Hosts or manages model deployments, agents, and tools.
  3. Azure API Management instance
    • Acts as the AI Gateway.
    • Receives requests from consumers.
    • Applies policies.
    • Routes requests to the appropriate backend.
  4. AI backend
    • Microsoft Foundry model.
    • Azure OpenAI deployment.
    • External model provider.
    • MCP server.
    • Other supported AI endpoint.
  5. Monitoring and governance services
    • API Management logs and metrics.
    • Application Insights, where configured.
    • Microsoft Foundry telemetry.
    • Security monitoring and auditing.

The gateway sits between the client and the AI backend. This allows the organization to enforce common controls without requiring every client application to implement those controls independently.


AI Gateway in Microsoft Foundry

Microsoft Foundry can be integrated with an Azure API Management instance as an AI Gateway.

This integration allows organizations to govern AI resources from within the Foundry environment while retaining access to the more advanced configuration capabilities of Azure API Management.

Depending on the supported feature and configuration, the gateway can help govern:

Models

The gateway can provide:

  • Token quotas.
  • Rate limits.
  • Authentication.
  • Routing.
  • Usage monitoring.
  • Model access control.
  • Centralized governance across model deployments.

When AI Gateway is used with Foundry, model requests can be routed through the associated API Management instance. Model limits can be configured at the project level, helping prevent one project or team from consuming all available capacity.

Agents

Agents can be registered and governed through Microsoft Foundry. Governance can include:

  • Centralized inventory.
  • Traffic policies.
  • Throttling.
  • Content safety controls.
  • Monitoring.
  • Access management.

The exact capabilities depend on the agent type and the integration being used.

Tools

MCP tools can be routed through an AI Gateway so that requests pass through a controlled endpoint.

Policies can be applied to MCP traffic, including:

  • Authentication.
  • Rate limiting.
  • IP filtering.
  • Correlation IDs.
  • Logging and metrics.
  • Routing controls.

However, the Foundry MCP gateway integration has limitations. For example, only eligible MCP tools created after the gateway is connected may be routed through the gateway. Existing tools are not automatically changed to use the gateway.


Prerequisites

Before configuring an AI Gateway for Microsoft Foundry, verify the following.

Azure API Management instance

You need an Azure API Management instance that meets the requirements for the selected integration.

Supported service tiers and networking requirements can vary depending on:

  • Whether the gateway is public or private.
  • Whether the Foundry resource has public network access disabled.
  • Whether private endpoints are required.
  • Whether advanced networking is needed.
  • Whether the organization is using the dedicated AI Gateway tier preview.

The dedicated AI Gateway tier is a public preview feature. Preview capabilities, supported regions, limits, and service behavior may change. Organizations should validate preview features carefully before using them for critical production workloads.

Required permissions

The administrator configuring the integration generally needs permission to manage the API Management instance.

For some Foundry gateway scenarios, the required role is:

  • API Management Service Contributor, or
  • Owner

The exact permissions depend on whether the administrator is connecting an existing gateway, creating an instance, importing models, or managing policies.

Networking

If the Foundry resource has public network access disabled, the API Management instance must also be able to access the private Foundry resource.

Depending on the architecture, this may require:

  • A private endpoint.
  • A supported API Management tier.
  • Virtual network integration or injection.
  • Appropriate private DNS configuration.
  • Network rules that permit the required traffic.

A gateway cannot securely mediate traffic to a private backend if the gateway itself cannot reach that backend.

Backend access

The administrator must have access to the model, deployment, or tool backend being added.

For managed identity authentication, the managed identity must also have the required permissions on the backend resource.


Creating or Associating an AI Gateway

The exact portal experience can change, but the general process is:

  1. Sign in to Microsoft Foundry.
  2. Open the appropriate Foundry administration or resource configuration area.
  3. Open the AI Gateway configuration.
  4. Select Add AI Gateway.
  5. Select the Foundry resource to associate with the gateway.
  6. Select an existing API Management instance or create one if supported.
  7. Confirm the required permissions and networking configuration.
  8. Save the association.
  9. Add or import models, agents, or eligible tools.
  10. Configure API Management policies.
  11. Test requests through the gateway.
  12. Verify telemetry and policy enforcement.

The Microsoft Foundry portal provides an integrated configuration experience, while advanced policies and networking settings are managed in Azure API Management.


Importing Models into the AI Gateway

Azure API Management can import models from supported providers, including Microsoft Foundry and Azure OpenAI.

When importing a model, the administrator typically configures:

  • The model provider.
  • The backend endpoint.
  • The model or deployment name.
  • The API format.
  • Authentication.
  • Required headers.
  • Backend routing.
  • Policies.
  • Monitoring settings.

For Microsoft Foundry deployments, the import wizard can discover deployments automatically in supported scenarios.

OpenAI-compatible APIs

Many applications are designed to use the OpenAI API format. API Management can expose supported backends through OpenAI-compatible routes.

For example, an OpenAI-compatible model endpoint may use a route similar to:

/default/models/openai/v1/chat/completions

The exact gateway URL and route depend on the API Management configuration and the provider API format.

The important exam concept is that the client can use a consistent gateway endpoint while API Management handles the connection to the underlying model backend.


Authentication Options

Authentication must be configured separately for:

  1. The client calling the gateway.
  2. The gateway calling the backend.

These are not necessarily the same authentication mechanism.

Client-to-gateway authentication

The client may authenticate to API Management using:

  • An API Management subscription key.
  • Microsoft Entra authentication.
  • OAuth.
  • Another supported API authentication method.

For example, a client may send an API Management subscription key in a header expected by the API definition or policy.

Gateway-to-backend authentication

The gateway can authenticate to AI backends using:

  • Managed identity.
  • API keys.
  • Provider-specific credentials.
  • Other supported authentication mechanisms.

Managed identity is often preferable because it avoids embedding long-lived API keys in applications or configuration files. The managed identity must have the appropriate role or permissions on the backend.

Managed identity benefits

Using managed identity can:

  • Avoid storing API keys in application code.
  • Reduce credential rotation requirements.
  • Integrate with Microsoft Entra access control.
  • Support centralized identity governance.
  • Reduce the risk of accidental credential exposure.

However, managed identity does not automatically grant access. The identity must still be authorized on the target resource.


API Management Policies

API Management policies are XML-based rules that execute in the gateway.

Policies can operate on:

  • Inbound requests.
  • Backend requests.
  • Backend responses.
  • Outbound responses.
  • Errors.

They can validate, transform, secure, route, or limit API traffic. API Management policies are different from Azure Policy. API Management policies run at API request time, while Azure Policy evaluates and governs Azure resources.

Important AI Gateway policies

Rate limiting

Rate limiting restricts the number of requests a client can make during a defined period.

Example uses include:

  • Limiting requests per application.
  • Limiting requests per user.
  • Preventing excessive tool calls.
  • Protecting backend capacity.
  • Reducing abuse.

A rate-limit-by-key policy can use a key such as:

  • Client IP address.
  • Subscription key.
  • Application identifier.
  • User identifier.
  • Custom request value.

The key should be selected carefully. IP-based limiting may be inappropriate when many users share the same outbound address.

Token quotas

AI model consumption is often measured in tokens rather than only requests.

Token quotas can help control:

  • Cost.
  • Capacity.
  • Fair usage.
  • Project-level consumption.
  • Large prompt abuse.
  • Excessive response generation.

A request-per-minute limit alone may not prevent a client from sending extremely large prompts. Token-based controls are therefore important for AI workloads.

IP filtering

IP filtering can restrict requests to trusted networks or addresses.

For example, an organization may allow gateway access only from:

  • Corporate networks.
  • Private application subnets.
  • Approved build environments.
  • Trusted automation services.

IP filtering should not be treated as a replacement for identity-based authentication. Network location alone is not sufficient to establish who is authorized to use an AI service.

Authentication and authorization

Policies can validate tokens, inspect claims, and enforce access rules.

For example, a policy may:

  • Validate a Microsoft Entra token.
  • Check the token audience.
  • Restrict access to specific application IDs.
  • Require a subscription key.
  • Reject unauthenticated requests.
  • Route different consumers to different backends.

Content safety

API Management can apply policies that integrate with Azure AI Content Safety to moderate prompts or responses.

Content safety controls may help detect or block content such as:

  • Hate.
  • Violence.
  • Sexual content.
  • Self-harm content.
  • Other unsafe material, depending on the configured policy and service capabilities.

Content safety is not a complete AI security solution. It should be combined with identity, authorization, data protection, logging, and runtime controls.

Correlation IDs

A correlation ID allows related requests to be traced across systems.

A gateway can add a unique identifier to a request so that administrators can correlate:

  • Client requests.
  • Gateway logs.
  • Backend requests.
  • Application logs.
  • Security investigations.

Correlation IDs are particularly useful when an application invokes multiple models or tools during one user interaction.

Request and response transformation

Policies can modify requests or responses, including:

  • Headers.
  • URLs.
  • Query parameters.
  • Payloads.
  • Backend routing.
  • Response formatting.

Transformations should be used carefully with AI APIs because changing required headers or payload structures can cause model or tool calls to fail.


Governing MCP Tools

The Model Context Protocol allows agents to interact with external tools and data sources.

Examples of MCP tools include:

  • Search tools.
  • File access tools.
  • Database tools.
  • Business application tools.
  • Automation tools.
  • Custom enterprise tools.

Routing MCP traffic through an AI Gateway provides a centralized point for:

  • Authentication.
  • Rate limiting.
  • IP restrictions.
  • Audit logging.
  • Routing.
  • Policy enforcement.

The gateway can apply controls without requiring changes to the MCP server or agent code.

Important MCP limitations

For the Foundry-integrated MCP gateway experience:

  • The feature is in preview.
  • Only eligible MCP tools are routed through the gateway.
  • Existing tools may not be automatically migrated.
  • Tools using managed OAuth may not be eligible for the same routing flow.
  • API Management policies are configured in Azure API Management.
  • Gateway logs may not contain complete tool-level traces.
  • MCP server logs may still be required for detailed tool investigation.

If a tool was created before the gateway was connected, it may continue to call the MCP server directly. In that situation, recreate the tool after the gateway is connected if the tool is eligible for gateway routing.


Monitoring and Observability

An AI Gateway provides a central location for monitoring AI traffic.

Useful telemetry may include:

  • Request counts.
  • Response codes.
  • Latency.
  • Backend failures.
  • Token usage.
  • Rate-limit events.
  • Authentication failures.
  • Policy violations.
  • Correlation IDs.
  • Model or backend usage.
  • Gateway errors.

Telemetry can be viewed through API Management monitoring capabilities and, where configured, Microsoft Foundry or Application Insights.

Monitoring helps answer questions such as:

  • Which applications are using a model?
  • Which projects consume the most tokens?
  • Are requests being rejected?
  • Is a backend unavailable?
  • Are clients exceeding quotas?
  • Are unusual traffic patterns occurring?
  • Are policies blocking legitimate workloads?
  • Are sensitive tools being called unexpectedly?

Gateway telemetry should be combined with application, model, agent, and backend logs because the gateway may not capture every detail of an AI interaction.


Security Design Considerations

Use a single governed entry point

Where practical, route approved AI traffic through the gateway rather than allowing every application to call model endpoints directly.

This improves consistency and visibility.

Avoid exposing backend credentials

Prefer managed identity or centrally managed credentials rather than embedding API keys in application code.

Apply least privilege

The gateway’s identity should have only the permissions required to access the configured backend.

The client should also receive only the permissions required to call the gateway APIs.

Separate environments

Use separate configurations or gateways for:

  • Development.
  • Testing.
  • Staging.
  • Production.

This helps prevent development applications from accessing production models or sensitive tools.

Apply quotas by project or consumer

Quotas should reflect business requirements. A shared global quota may allow one application to consume capacity needed by other teams.

Restrict sensitive tools

Tools that can:

  • Modify databases.
  • Send email.
  • Change permissions.
  • Deploy resources.
  • Access confidential data.
  • Execute commands.

should receive stronger controls than read-only tools.

Protect private backends

If a Foundry resource is private, ensure the gateway has appropriate private connectivity. Do not assume that associating the resources automatically solves networking requirements.

Test policies before enforcement

Use testing and staged rollout to ensure that policies do not:

  • Remove required authentication headers.
  • Break model payloads.
  • Block legitimate clients.
  • Prevent required tool calls.
  • Cause unexpected latency.
  • Interfere with streaming responses.

Example Scenario

A company has three AI applications:

  • A customer-service assistant.
  • A financial reporting application.
  • An internal research assistant.

Each application currently calls a model endpoint directly.

The security team deploys Azure API Management as an AI Gateway and configures:

  1. Microsoft Entra authentication for approved applications.
  2. Managed identity authentication from the gateway to the model backend.
  3. Separate API products or subscriptions for each application.
  4. Token quotas for each project.
  5. Rate limits to prevent excessive requests.
  6. IP restrictions for internal applications.
  7. Content safety checks.
  8. Correlation IDs for tracing.
  9. Monitoring and alerts for failures and abnormal usage.

The financial reporting application receives a lower quota but stronger access restrictions because it processes sensitive information. The research assistant is allowed to use several approved models but cannot access production business tools.

This design provides centralized security while allowing different applications to have different access and usage policies.


Common Exam Comparisons

CapabilityPrimary purpose
Azure API Management AI GatewayGovern and secure traffic to AI models, agents, and tools
Microsoft FoundryDevelop, deploy, and operate AI resources
Microsoft Entra IDAuthenticate identities and authorize access
Managed identityProvide Azure-managed authentication without embedded credentials
API Management policyApply runtime request and response controls
Azure PolicyGovern Azure resource configuration and compliance
Azure AI Content SafetyDetect or moderate unsafe content
Application InsightsMonitor application and gateway-related telemetry
Microsoft SentinelCentralize security events and investigation workflows
Agent identityRepresent an AI agent as an identity
MCP serverExpose tools or context to compatible AI clients

Exam Tips

Remember these key points:

  • An AI Gateway is implemented using Azure API Management capabilities.
  • The gateway is a control point between AI consumers and AI backends.
  • It can govern models, agents, and eligible MCP tools.
  • The gateway is not a replacement for Microsoft Entra authentication.
  • Client-to-gateway authentication and gateway-to-backend authentication are separate concerns.
  • Managed identity can reduce the need to store backend API keys.
  • API Management policies are runtime rules, not Azure Policy definitions.
  • Rate limits control request volume; token quotas control AI consumption.
  • Content safety policies address unsafe content but do not replace identity security.
  • Existing MCP tools may not automatically begin using a newly connected gateway.
  • API Management policies are configured in API Management, not necessarily in the Foundry portal.
  • Private Foundry resources require compatible private connectivity from the gateway.
  • Preview features may have changing capabilities, limits, and supported regions.
  • Monitoring should include gateway telemetry and backend or application logs.
  • Always verify that traffic actually passes through the gateway after configuration.

Practice Exam Questions

Question 1

What is the primary purpose of using Azure API Management as an AI Gateway for Microsoft Foundry?

A. To replace Microsoft Entra ID
B. To train foundation models
C. To provide a centralized point for securing, governing, and monitoring AI traffic
D. To permanently store model training data

Answer: C

Explanation: An AI Gateway provides a centralized control point between AI consumers and backends. It can enforce authentication, quotas, rate limits, routing, monitoring, and other policies. It does not replace Microsoft Entra ID or train models.


Question 2

An organization wants the AI Gateway to authenticate to a Microsoft Foundry backend without storing an API key in application code. Which option should it consider?

A. Managed identity
B. Azure resource lock
C. Public IP address filtering only
D. A storage account access key

Answer: A

Explanation: Managed identity allows the gateway to authenticate to supported Azure services without embedding long-lived credentials in application code. The identity must still be granted the required permissions on the backend.


Question 3

Which statement correctly describes Azure API Management policies?

A. They are Azure resource-compliance definitions evaluated by Azure Policy
B. They are runtime rules that can validate, transform, secure, limit, or route API requests and responses
C. They are Microsoft Entra role assignments
D. They are model-training instructions

Answer: B

Explanation: API Management policies execute in the gateway and can control inbound requests, backend requests, responses, and errors. They are different from Azure Policy, which governs Azure resource configuration.


Question 4

A company wants to prevent one Foundry project from consuming all available model capacity. Which control is most directly relevant?

A. Azure Bastion
B. Microsoft Entra access reviews
C. Token quotas
D. Azure resource locks

Answer: C

Explanation: Token quotas limit AI consumption based on token usage. They are useful for cost control, capacity management, and preventing one project from monopolizing model capacity.


Question 5

An administrator connects an AI Gateway to a Foundry resource. An MCP tool created several weeks earlier continues to call the MCP server directly. What is the most likely explanation?

A. API Management cannot govern MCP tools
B. The tool must be recreated after the gateway is connected if it is eligible for gateway routing
C. The Foundry resource must be deleted
D. The MCP server must be converted into a virtual machine

Answer: B

Explanation: In the preview Foundry MCP gateway integration, gateway routing is applied when eligible tools are created after the gateway is connected. Existing tools are not automatically migrated.


Question 6

Which control is most appropriate for restricting AI Gateway requests to approved corporate networks?

A. IP filtering
B. Token quota
C. Model fine-tuning
D. Agent blueprint inheritance

Answer: A

Explanation: IP filtering can allow or deny requests based on source IP addresses or ranges. It should supplement, not replace, identity-based authentication and authorization.


Question 7

An organization wants to trace a request from an application through API Management to the model backend. Which feature is most useful?

A. Azure resource locks
B. Correlation IDs
C. Disk encryption
D. Azure Policy remediation

Answer: B

Explanation: Correlation IDs provide a common identifier that can be recorded in gateway, application, and backend logs, making it easier to trace a request across multiple components.


Question 8

A Foundry resource has public network access disabled. What must be verified before associating it with an AI Gateway?

A. The gateway has an appropriate private network path to the Foundry resource
B. The model has been fine-tuned
C. All clients use anonymous access
D. The gateway is deployed outside Azure

Answer: A

Explanation: A private Foundry resource requires compatible private connectivity from the API Management instance. The gateway must be able to reach the backend through the required private networking configuration.


Question 9

Which statement best describes the difference between rate limiting and token quotas?

A. Rate limiting controls request volume, while token quotas control AI token consumption
B. Rate limiting authenticates users, while token quotas encrypt traffic
C. Rate limiting governs Azure resources, while token quotas manage virtual machines
D. Rate limiting disables models, while token quotas create agents

Answer: A

Explanation: Rate limiting restricts the number of requests over a period. Token quotas restrict the amount of model consumption, which is important because requests can vary greatly in prompt and response size.


Question 10

An administrator applies an API Management policy that removes a required authentication header before forwarding the request to an MCP server. What is the likely result?

A. The MCP server automatically repairs the header
B. The request is converted into a model-training job
C. The request may fail because required authentication information was removed
D. The gateway automatically grants anonymous access

Answer: C

Explanation: API Management policies can modify headers, but removing a required authentication header can cause the backend or MCP server to reject the request. Policies must be tested carefully to avoid breaking required authentication and protocol behavior.


Final Exam Point

A very important exam theme is that Azure API Management provides the enforcement and governance layer, while Microsoft Foundry provides the AI resource and application environment. Together, they allow organizations to centralize access control, usage management, monitoring, and security for AI workloads.


Go to the SC-500 Exam Prep Hub main page