> ## Documentation Index
> Fetch the complete documentation index at: https://bifrost-dev.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Ask Warp a question

> Runs Warp's agent loop over the deployment's telemetry and returns an
answer.

**The route is always registered, and answers `503` while Warp cannot
work.** The body carries a machine-readable `reason`, so "present but
unusable" stays distinguishable from "absent" without the state being
permanent: a deployment that enables logging after startup begins serving
this route without a restart.

The model's context is client-sent: `messages` carries the thread so far.
Where Warp history is available the server also files each turn under
`conversation_id` so the thread can be listed and reopened through
`/api/warp/conversations`. Omit `conversation_id` to start a new thread;
its id comes back on the `done` event or in the JSON body.

**Data visibility equals the caller's own.** Every tool query runs against
the caller's row-level scope, so Warp can never surface data the caller
could not already retrieve from the logs API directly.

The turn runs on Warp's default model unless the body names another with
`provider` and `model`. Only a model the configuration exposes is
accepted; anything else is a `400` and no model is called.

With `stream: true` (the default) the response is `text/event-stream`. Both
modes run the identical loop; only the sink differs.

Enterprise RBAC: requires `WarpSession` View.




## OpenAPI

````yaml /openapi/openapi.json post /api/warp/chat
openapi: 3.1.0
info:
  title: Bifrost API
  description: >
    Bifrost HTTP Transport API for AI model inference and gateway management.


    This API provides a unified interface for interacting with multiple AI
    providers

    including OpenAI, Anthropic, Bedrock, Gemini, and more through a single API,

    along with comprehensive management APIs for configuring and monitoring the
    gateway.


    ## API Structure


    ### Unified Inference API (`/v1/*`)

    The primary API using Bifrost's unified format. Model parameters use the
    format

    `provider/model` (e.g., `openai/gpt-4`, `anthropic/claude-3-opus`).


    ### Async Inference API (`/v1/async/*`)

    Submit inference requests for asynchronous execution. Returns a job ID
    immediately

    and allows polling for results. Supports all inference types except batches,
    files,

    and containers.


    ### Provider Integration APIs

    Native provider-format APIs for drop-in compatibility:

    - `/openai/*` - OpenAI-compatible API

    - `/anthropic/*` - Anthropic-compatible API

    - `/genai/*` - Google GenAI (Gemini) compatible API

    - `/bedrock/*` - AWS Bedrock compatible API

    - `/cohere/*` - Cohere compatible API


    ### Framework Integration APIs

    Multi-provider proxy endpoints for AI frameworks:

    - `/litellm/*` - LiteLLM proxy with all provider formats

    - `/langchain/*` - LangChain compatible endpoints

    - `/pydanticai/*` - PydanticAI compatible endpoints


    ### Management APIs (`/api/*`)

    APIs for managing and monitoring the Bifrost gateway:

    - `/api/config` - Configuration management

    - `/api/providers` - Provider and API key management

    - `/api/plugins` - Plugin management

    - `/api/governance/*` - Virtual keys, teams, customers, budgets, rate
    limits, routing rules, and pricing overrides

    - `/api/logs` - Log search and analytics

    - `/api/mcp/*` - MCP (Model Context Protocol) client management

    - `/api/session/*` - Authentication and session management

    - `/api/cache/*` - Cache management

    - `/health` - Health check endpoint


    ## Fallbacks

    Requests can include fallback models that will be tried if the primary model
    fails.
  version: 1.0.0
  contact:
    name: Contact Us
    url: https://getmaxim.ai/bifrost
  license:
    name: Apache 2.0
    url: https://opensource.org/licenses/Apache-2.0
servers:
  - url: '{baseUrl}'
    description: Your Bifrost instance
    variables:
      baseUrl:
        default: http://localhost:8080
        description: Base URL of your Bifrost instance (e.g. https://bifrost.mycompany.com)
security:
  - BearerAuth: []
  - BasicAuth: []
  - ApiKeyAuth: []
tags:
  - name: Models
    description: Model listing and information
  - name: Chat Completions
    description: Chat-based text generation
  - name: Text Completions
    description: Text completion generation
  - name: Responses
    description: OpenAI Responses API compatible endpoints
  - name: OCR
    description: Optical character recognition for documents and images
  - name: Rerank
    description: Document reranking by relevance to a query
  - name: Decisions
    description: Structured decisions evaluated against annotated function-tool definitions
  - name: Embeddings
    description: Text embedding generation
  - name: Images
    description: Image generations, editing, and variations
  - name: Videos
    description: Video generation and management
  - name: Audio
    description: Speech synthesis and transcription
  - name: Count Tokens
    description: Token counting utilities
  - name: Batch
    description: Batch processing operations
  - name: Files
    description: File management operations
  - name: Containers
    description: Container management operations
  - name: Async Jobs
    description: Asynchronous job submission and retrieval endpoints
  - name: Realtime
    description: Realtime WebSocket and WebRTC endpoints
  - name: OpenAI Integration
    description: OpenAI-compatible API endpoints (/openai/*)
  - name: Azure Integration
    description: Azure OpenAI integration endpoints
  - name: Anthropic Integration
    description: Anthropic-compatible API endpoints (/anthropic/*)
  - name: GenAI Integration
    description: Google GenAI (Gemini) compatible API endpoints (/genai/*)
  - name: Bedrock Integration
    description: AWS Bedrock compatible API endpoints (/bedrock/*)
  - name: Cohere Integration
    description: Cohere compatible API endpoints (/cohere/*)
  - name: Typesafe Integration
    description: Typesafe compatible API endpoints (/typesafe/*)
  - name: LiteLLM Integration
    description: LiteLLM proxy endpoints with multi-provider support (/litellm/*)
  - name: LangChain Integration
    description: LangChain compatible endpoints with multi-provider support (/langchain/*)
  - name: PydanticAI Integration
    description: >-
      PydanticAI compatible endpoints with multi-provider support
      (/pydanticai/*)
  - name: Health
    description: Health check endpoints
  - name: Configuration
    description: Configuration management endpoints
  - name: Session
    description: Session and authentication endpoints
  - name: Providers
    description: Provider management endpoints
  - name: Plugins
    description: Plugin management endpoints
  - name: MCP
    description: Model Context Protocol endpoints
  - name: Governance
    description: Virtual keys, teams, and customers management
  - name: Routing
    description: Routing rules and complexity analyzer configuration
  - name: Logging
    description: Log search and management endpoints
  - name: Cache
    description: Cache management endpoints
  - name: Vault
    description: Vault secret management endpoints
  - name: Skills
    description: Skills Repository management, marketplace, and download endpoints
  - name: Audit Logs
    description: >-
      CADF-compliant audit log search, export, and signature verification
      endpoints
  - name: Webhooks
    description: Webhook endpoint management and signed async-job delivery history
  - name: Notifications
    description: >-
      Role-targeted dashboard notifications, delivered over the dashboard
      WebSocket
  - name: Background Jobs
    description: >-
      Status, progress and cancellation of durable background jobs (cost
      recalculation, Warp log indexing, and others)
  - name: Warp
    description: >-
      Configuration for Warp, the dashboard agent that answers questions about
      the deployment's own telemetry
paths:
  /api/warp/chat:
    post:
      tags:
        - Warp
      summary: Ask Warp a question
      description: >
        Runs Warp's agent loop over the deployment's telemetry and returns an

        answer.


        **The route is always registered, and answers `503` while Warp cannot

        work.** The body carries a machine-readable `reason`, so "present but

        unusable" stays distinguishable from "absent" without the state being

        permanent: a deployment that enables logging after startup begins
        serving

        this route without a restart.


        The model's context is client-sent: `messages` carries the thread so
        far.

        Where Warp history is available the server also files each turn under

        `conversation_id` so the thread can be listed and reopened through

        `/api/warp/conversations`. Omit `conversation_id` to start a new thread;

        its id comes back on the `done` event or in the JSON body.


        **Data visibility equals the caller's own.** Every tool query runs
        against

        the caller's row-level scope, so Warp can never surface data the caller

        could not already retrieve from the logs API directly.


        The turn runs on Warp's default model unless the body names another with

        `provider` and `model`. Only a model the configuration exposes is

        accepted; anything else is a `400` and no model is called.


        With `stream: true` (the default) the response is `text/event-stream`.
        Both

        modes run the identical loop; only the sink differs.


        Enterprise RBAC: requires `WarpSession` View.
      operationId: warpSession
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/WarpSessionRequest'
      responses:
        '200':
          description: >
            The answer. `text/event-stream` when streaming, otherwise a single
            JSON

            body.


            SSE frames are `event: <type>` with a JSON payload. Types are
            `start`,

            `delta` (an answer fragment), `tool_call_start`, `tool_call_end`,

            `error` and `done`.


            **`error` is terminal and is never followed by `done`.** A client
            that

            treats `done` as the only completion signal will read a failed
            request

            as a successful one, so both must end the stream.
          content:
            text/event-stream:
              schema:
                type: string
            application/json:
              schema:
                $ref: '#/components/schemas/WarpSessionResponse'
        '400':
          description: Bad request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BifrostError'
        '403':
          description: Caller lacks `WarpSession` View (enterprise)
        '413':
          description: Conversation exceeds the server's history limit
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BifrostError'
        '500':
          description: Internal server error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BifrostError'
        '503':
          description: |
            Warp cannot answer yet. `reason` says why: `no_log_store` when the
            deployment persists no logs, `not_configured` when it has no usable
            settings or no model client. The dashboard treats them differently -
            the first hides the launcher, the second links to settings.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/WarpUnavailable'
      security:
        - ManagementBearerAuth: []
components:
  schemas:
    WarpSessionRequest:
      type: object
      properties:
        messages:
          type: array
          minItems: 1
          description: >
            The full conversation so far, oldest first. Roles must be `user` or

            `assistant`; a client-supplied `system` turn is rejected, since it
            would

            let a caller displace the instructions that keep Warp from inventing

            numbers.


            Long threads are trimmed server-side, always keeping the opening
            turn -

            it usually carries the framing everything after it depends on.
          items:
            type: object
            properties:
              role:
                type: string
                enum:
                  - user
                  - assistant
              content:
                type: string
            required:
              - role
              - content
        conversation_id:
          type: string
          maxLength: 36
          description: >
            Continues an existing thread in the caller's history. Omit it to
            start a

            new one; the new thread's id comes back on the `done` event or in
            the

            JSON body.
        provider:
          type: string
          description: >
            With `model`, names which of Warp's configured models answers this
            turn:

            the default, or one of `additional_models`. Omit both to use the

            default. A pair the operator has not exposed is rejected with `400`,
            so

            a request can only ever choose among the configured models, and the

            provider key is always the configured entry's own.
          example: anthropic
        model:
          type: string
          description: See `provider`. The two are given together or not at all.
          example: claude-sonnet-5
        stream:
          type: boolean
          default: true
          description: |
            Selects the transport, not the behaviour. `true` streams SSE frames;
            `false` returns one JSON body. Both run the same loop.
      required:
        - messages
    WarpSessionResponse:
      type: object
      description: The non-streaming body - the same events, assembled.
      properties:
        answer:
          type: string
        tool_calls:
          type: array
          description: |
            What Warp queried, in order. The result payloads are deliberately
            absent: the model consumed them, and echoing them would multiply the
            response size for no reader benefit.
          items:
            type: object
            properties:
              name:
                type: string
              arguments:
                type: string
              duration_ms:
                type: integer
              failed:
                type: boolean
            required:
              - name
              - duration_ms
        iterations:
          type: integer
          description: How many model round trips the answer took.
        conversation_id:
          type: string
          description: >-
            The thread this turn was filed under. Absent when Warp history is
            unavailable.
        finish_reason:
          type: string
        usage:
          allOf:
            - $ref: '#/components/schemas/BifrostLLMUsage'
          description: >
            Token usage summed across every model round trip in the request.
            Warp

            runs on its own Bifrost instance with no logging plugin -
            deliberately,

            so its calls do not pollute the log table it reads - which means
            this is

            the only place its spend is reported.
        error:
          type: object
          description: >
            Present when the request failed. The codes are the closed set the
            agent

            emits; both fields are always populated when the object is present.
          properties:
            code:
              type: string
              enum:
                - upstream_error
                - access_denied
                - max_iterations
                - timeout
                - cancelled
            message:
              type: string
          required:
            - code
            - message
      required:
        - answer
        - tool_calls
        - iterations
    BifrostError:
      type: object
      description: Error response from Bifrost
      properties:
        event_id:
          type: string
        type:
          type: string
        is_bifrost_error:
          type: boolean
        status_code:
          type: integer
        error:
          $ref: '#/components/schemas/ErrorField'
        extra_fields:
          $ref: '#/components/schemas/BifrostErrorExtraFields'
    WarpUnavailable:
      type: object
      description: >
        The 503 body. `reason` exists because the causes need different
        treatment in

        the dashboard, and a client would otherwise have to pattern-match the

        message, which breaks silently the first time it is reworded.


        - `not_configured`: fixable by the operator. Stay visible and link to
        Warp's
          settings page.
        - `no_log_store`: the deployment persists no logs, so Warp has nothing
        to
          read and there is no in-panel remedy. Hide the launcher.
        - `no_vector_store`: logs exist but no vector store is connected, so
          semantic search specifically is unavailable. Everything else Warp does
          still works, so this is a degraded state rather than an absent one - do
          not hide the launcher for it.
      properties:
        reason:
          type: string
          enum:
            - not_configured
            - no_log_store
            - no_vector_store
        message:
          type: string
      required:
        - reason
        - message
    BifrostLLMUsage:
      type: object
      description: Token usage information
      properties:
        prompt_tokens:
          type: integer
          description: >
            Total input tokens including any prompt-cache tokens (read + write).
            Subtract prompt_tokens_details.cached_read_tokens and
            prompt_tokens_details.cached_write_tokens to get the non-cached
            portion.
        prompt_tokens_details:
          $ref: '#/components/schemas/ChatPromptTokensDetails'
        completion_tokens:
          type: integer
          description: Number of output/completion tokens generated.
        completion_tokens_details:
          $ref: '#/components/schemas/ChatCompletionTokensDetails'
        total_tokens:
          type: integer
        tool_usage:
          type: object
          description: |
            Counts of provider-hosted (server-side) tool calls billed per call.
          properties:
            web_search:
              type: object
              properties:
                num_requests:
                  type: integer
                  description: Number of billable web search calls.
        cost:
          $ref: '#/components/schemas/BifrostCost'
    ErrorField:
      type: object
      properties:
        type:
          type: string
        code:
          type: string
        message:
          type: string
        param:
          type: string
        event_id:
          type: string
    BifrostErrorExtraFields:
      type: object
      properties:
        provider:
          $ref: '#/components/schemas/ModelProvider'
        model_requested:
          type: string
        request_type:
          type: string
        error_type:
          type: string
          description: >-
            Normalized, low-cardinality classification of why the request
            failed, declared by whichever component refused it. Prefixed by
            fault domain (caller_, policy_, provider_, bifrost_). Absent on
            failures that were not classified at source.
        retry_after_ms:
          type: integer
          format: int64
          minimum: 1000
          maximum: 300000
          description: >-
            The provider's hint for how long to wait before retrying, in
            milliseconds, read from its retry-after-ms or Retry-After header or
            its google.rpc.RetryInfo error detail, and clamped to between 1000
            and 300000. Absent when the provider gave no explicit hint.
    ChatPromptTokensDetails:
      type: object
      properties:
        text_tokens:
          type: integer
        audio_tokens:
          type: integer
        image_tokens:
          type: integer
        cached_read_tokens:
          type: integer
          description: >
            Tokens served from the prompt cache (cache hit). These tokens are
            already included in prompt_tokens and are billed at the reduced
            cache-read rate. Populated for all providers that support prompt
            caching (Anthropic, Bedrock, OpenAI, Gemini, xAI, etc.).
        cached_write_tokens:
          type: integer
          description: >
            Tokens written to the prompt cache on this request (cache creation /
            write). These tokens are already included in prompt_tokens and are
            billed at the cache-creation rate. Populated for providers that
            separately report cache write tokens (Anthropic, Bedrock).
    ChatCompletionTokensDetails:
      type: object
      properties:
        text_tokens:
          type: integer
        accepted_prediction_tokens:
          type: integer
        audio_tokens:
          type: integer
        citation_tokens:
          type: integer
        num_search_queries:
          type: integer
          deprecated: true
          description: >-
            Deprecated, use tool_usage.web_search.num_requests. Will be removed
            in 3.0.0.
        reasoning_tokens:
          type: integer
        image_tokens:
          type: integer
        rejected_prediction_tokens:
          type: integer
    BifrostCost:
      type: object
      description: Cost breakdown for the request
      properties:
        input_tokens_cost:
          type: number
        output_tokens_cost:
          type: number
        reasoning_tokens_cost:
          type: number
          description: Cost for reasoning/thinking tokens (reasoning models)
        citation_tokens_cost:
          type: number
          description: Cost for citation tokens
        search_queries_cost:
          type: number
          description: Cost for web search queries
        request_cost:
          type: number
        total_cost:
          type: number
    ModelProvider:
      type: string
      description: AI model provider identifier
      enum:
        - anthropic
        - azure
        - bedrock
        - bedrock_mantle
        - cerebras
        - cohere
        - deepseek
        - gemini
        - groq
        - mistral
        - ollama
        - opencode-go
        - opencode-zen
        - openai
        - parasail
        - perplexity
        - sgl
        - vertex
        - openrouter
        - elevenlabs
        - huggingface
        - nebius
        - xai
        - replicate
        - vllm
        - runway
        - runware
        - fireworks
        - sarvam
        - wafer
        - databricks
        - typesafe
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      description: >
        Bearer token authentication. Use your provider API key or Bifrost
        authentication token.

        Virtual keys (prefixed with `sk-bf-`) can also be passed here.
    BasicAuth:
      type: http
      scheme: basic
      description: >
        Basic authentication using the Bifrost admin username and password

        (`auth_config.admin_username` / `auth_config.admin_password`).

        Accepted on management APIs (`/api/*`, `/metrics`, `/ws`) only - the
        inference

        middleware never validates Basic credentials.
    ApiKeyAuth:
      type: apiKey
      in: header
      name: x-api-key
      description: |
        API key authentication via the `x-api-key` header.
        Virtual keys (prefixed with `sk-bf-`) can also be passed here.
    ManagementBearerAuth:
      type: http
      scheme: bearer
      description: >
        Management API authentication for `/api/*` endpoints. Use the
        `Authorization` header

        with `Bearer <token>`, where `<token>` is one of:


        - a Bifrost management API key,

        - a dashboard session token issued by `POST /api/session/login`,

        - base64 of `<admin-username>:<admin-password>` (legacy equivalent of
        `BasicAuth`).


        Virtual keys (`sk-bf-*`) and the `x-api-key` header are not accepted on
        management APIs -

        the sole exception is `GET /api/governance/virtual-keys/quota`, which is
        virtual-key-only.


        Authentication alone is not sufficient in Bifrost Enterprise: each
        operation page shows a

        **Required Permissions** table (`Resource:Operation`, for example
        `Dashboard:View`) above

        its Authorizations section, and the caller's RBAC role or management API
        key scopes must

        include what it lists, otherwise the request is rejected with `403
        Forbidden`.


        A local admin — authenticated with the admin password, or any caller on
        a deployment with

        dashboard auth disabled — bypasses these checks and can call every
        management endpoint.


        **OSS setup lock.** On Bifrost OSS, while dashboard auth is not active
        (no admin account,

        or auth disabled), every management endpoint except the public ones
        (`/health`,

        `/api/version`, `/api/session/is-auth-enabled`, `/api/session/login`,
        ...) requires the

        operator's setup token in the `X-Bifrost-Setup-Token` header, in place
        of `Authorization`.

        The token is set with `setup_token` in `config.json` or the
        `BIFROST_SETUP_TOKEN`

        environment variable. A missing header returns `401`, a wrong token
        `403`. The header

        stops working once dashboard auth is enabled. The dashboard instead
        trades the token once

        for an HttpOnly `bifrost_setup_session` cookie via `POST
        /api/session/setup`.

        See [Required permissions](/api/procuring-api-keys#required-permissions)
        for how

        permissions are derived and which endpoints are exempt.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.