NVIDIA Run:ai Distributed Inferences API

Distributed inference enables running inference workloads across multiple pods, typically to scale model serving beyond a single container or node. This approach is useful when a single instance cannot meet resource requirements.NVIDIA Run:ai supports this model using Leader Worker Set (LWS). Each pod plays a specific role, either as a leader or worker, and together they form a coordinated service. NVIDIA Run:ai manages the orchestration and configuration of these pods to ensure efficient and scalable inference execution

Operations 4

POST /api/v1/workloads/distributed-inferences Create a distributed inference. [Experimental] #
DELETE /api/v1/workloads/distributed-inferences/{workloadId} Delete a distributed inference. #
GET /api/v1/workloads/distributed-inferences/{workloadId} Get a distributed inference data. #
PATCH /api/v1/workloads/distributed-inferences/{workloadId} Update distributed inference spec. #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/runai-distributed-inferences-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

runai-distributed-inferences-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  version: latest
  description: '# Introduction


    The NVIDIA Run:ai Control-Plane API reference is a guide that provides an easy-to-use programming interface for adding various tasks to your application, including workload submission, resource management, and administrative operations.


    NVIDIA Run:ai APIs are accessed using *bearer tokens*. To obtain a token, you need to create a **Service account** through the NVIDIA Run:ai user interface.

    To create a service account, in your UI, go to Access → Service Accounts (for organization-level service accounts) or User settings → Access Keys (for user access keys), and create a new one.


    After you have created a new service account, you will need to assign it access rules.

    To assign access rules to the service account, see [Create access rules](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/accessrules#create-or-delete-rules).

    Make sure you assign the correct rules to your service account. Use the [Roles](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/roles) to assign the correct access rules.


    To get your access token, follow the instructions in [Request a token](https://run-ai-docs.nvidia.com/saas/reference/api/rest-auth/#request-an-api-token).

    '
  title: NVIDIA Run:ai Access Keys Distributed Inferences API
  x-logo:
    url: https://api.redocly.com/registry/raw/runai-xq8/saas/latest/public/runai-logo-api.png
    altText: NVIDIA Run:ai
    href: https://run.ai
  license:
    name: NVIDIA Run:ai
    url: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-license-agreement/
servers:
- url: https://app.run.ai
security:
- bearerAuth: []
tags:
- name: Distributed Inferences
  description: "Distributed inference enables running inference workloads across multiple pods, typically to scale model serving beyond a single container or node. This approach is useful when a single instance cannot meet resource requirements.NVIDIA Run:ai supports this model using Leader Worker Set (LWS). \nEach pod plays a specific role, either as a leader or worker, and together they form a coordinated service. NVIDIA Run:ai manages the orchestration and configuration of these pods to ensure efficient and scalable inference execution\n"
paths:
  /api/v1/workloads/distributed-inferences:
    post:
      summary: Create a distributed inference. [Experimental]
      operationId: create_distributed_inference
      description: Create a distributed inference using container related fields.
      tags:
      - Distributed Inferences
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/DistributedInferenceCreationRequest'
      responses:
        '202':
          description: Request completed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DistributedInference'
        '400':
          $ref: '#/components/responses/400BadRequest'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
  /api/v1/workloads/distributed-inferences/{workloadId}:
    delete:
      summary: Delete a distributed inference.
      operationId: delete_distributed_inference
      description: Delete a distributed inference using a workload id.
      tags:
      - Distributed Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      responses:
        '202':
          $ref: '#/components/responses/202Accepted'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
    get:
      summary: Get a distributed inference data.
      operationId: get_distributed_inference
      description: Retrieve a distributed inference details using a workload id.
      tags:
      - Distributed Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      responses:
        '200':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DistributedInference'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
    patch:
      summary: Update distributed inference spec.
      operationId: update_distributed_inference_spec
      description: Update the specification of an existing distributed inference workload.
      tags:
      - Distributed Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/UpdateRequest'
      responses:
        '202':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DistributedInference'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
components:
  schemas:
    LargeShmRequest:
      description: A large /dev/shm device to mount into a container running the created workload. An shm is a shared file system mounted on RAM.
      type:
      - boolean
      - 'null'
      example: false
    SecretFieldsUpdatable:
      properties:
        mountPath:
          description: Local path within the workload to which the Secret will be mapped to. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
        defaultMode:
          $ref: '#/components/schemas/DefaultMode'
      type:
      - object
      - 'null'
    EnvironmentVariableConfigMap:
      description: Details of the configMap and key use to populate the environment variable
      properties:
        name:
          description: The name of the config-map resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: my-config-map
          pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
        key:
          description: The key in the config-map resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: MY_POSTGRES_SCHEMA
          pattern: .*
      type:
      - object
      - 'null'
    ConfigMapInstance:
      allOf:
      - $ref: '#/components/schemas/StorageInstanceName'
      - $ref: '#/components/schemas/ConfigMap'
      - $ref: '#/components/schemas/ExcludeField'
      type:
      - object
      - 'null'
    WorkloadId2:
      description: A unique ID of the workload.
      type: string
      format: uuid
    DefaultMode:
      type:
      - string
      - 'null'
      description: 'File permission mode in octal string format. This value must be a 4-digit octal number, representing the default file mode when mounting a Secret or ConfigMap as a volume.

        '
      minLength: 4
      maxLength: 4
      example: '0644'
      pattern: 0[0-7]{3}
    CpuMemoryRequest:
      description: The amount of CPU memory to allocate for this workload (1G, 20M, .etc). The workload will receive at least this amount of memory. Note that the workload will not be scheduled unless the system can guarantee this amount of memory to the workload
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 20M
    GpuDevicesRequest:
      description: Requested number of GPU devices. Currently if more than one device is requested, it is not possible to provide values for gpuMemory or gpuPortion.
      type:
      - integer
      - 'null'
      format: int32
      example: 1
      minimum: 0
    PriorityClass:
      description: 'Specifies the priority class for the workload, which determines its scheduling behavior. Valid values are: very-low, low, medium-low, medium, medium-high, high, and very-high. Each workload type has a default priority. To view the default priority for each workload type, use the GET /workload-types endpoint. Once you change the priority from the default value defined for that workload type, the preemptibility field is not automatically updated. Make sure to set the desired preemptibility value.'
      type:
      - string
      - 'null'
      pattern: .*
    Label:
      description: Label details to be populated into the container.
      properties:
        name:
          description: The name of the label (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: stage
          pattern: .*
        value:
          description: The value of the label.
          type:
          - string
          - 'null'
          example: initial-research
          pattern: .*
        exclude:
          description: Use 'true' in case the label is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
      type:
      - object
      - 'null'
    DepartmentId2:
      description: The id of the department.
      type: string
      minLength: 1
      example: 2
      pattern: .*
    DistributedInferenceStartupPolicyField:
      properties:
        startupPolicy:
          $ref: '#/components/schemas/DistributedInferenceStartupPolicy'
      type:
      - object
      - 'null'
    NodeSelectorTerm:
      type:
      - object
      - 'null'
      description: A null or empty node selector term matches no objects. The requirements of them are ANDed.
      properties:
        matchExpressions:
          description: A list of node selector requirements by node's labels.
          type: array
          items:
            $ref: '#/components/schemas/MatchExpression'
    ImagePullSecrets:
      description: A list of references to Kubernetes secrets in the same namespace used for pulling container images.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/ImagePullSecret'
    PvcVolumeMode:
      description: Default volume mode for the PVC. Choose between Filesystem (default) or Block.
      type:
      - string
      - 'null'
      enum:
      - Filesystem
      - Block
    TolerationEffect:
      description: The taint effect to match. (mandatory)
      type:
      - string
      - 'null'
      enum:
      - NoSchedule
      - NoExecute
      - PreferNoSchedule
      - Any
    Annotation:
      description: Annotation details to be populated into the container.
      properties:
        name:
          description: The name of the annotation (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: billing
          pattern: .*
        value:
          description: The value of the annotation.
          type:
          - string
          - 'null'
          example: my-billing-unit
          pattern: .*
        exclude:
          description: Use 'true' in case the annotation is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          default: false
          example: false
      type:
      - object
      - 'null'
    NodeType3:
      description: Nodes (machines), or a group of nodes on which the workload will run. To use this feature, your Administrator will need to label nodes. For more information, see [Group Nodes](https://docs.run.ai/latest/admin/researcher-setup/limit-to-node-group). When using this flag with with Project-based affinity, it refines the list of allowable node groups set in the Project. For more information, see [Projects](https://docshub.run.ai/guides/platform-management/aiinitiatives/organization/projects).
      type:
      - string
      - 'null'
      minLength: 1
      example: my-node-type
      pattern: .*
    PvcClaimSize:
      description: 'Requested size for the PVC. Mandatory when existingPvc is false. Recommended sizes: TB/GB/MB/TIB/GIB/MIB'
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 1G
    EnvironmentVariablePodFieldReference:
      description: Details of the field-reference and key use to populate the environment variable
      properties:
        path:
          description: The field path resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: metadata.name
          pattern: .*
      type:
      - object
      - 'null'
    Labels:
      description: Set of labels to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Label'
    PvcFieldsNonUpdatable:
      properties:
        existingPvc:
          description: Verify existing PVC. PVC is assumed to exist when set to `true`. If set to `false`, the PVC will be created, if it does not exist.
          type:
          - boolean
          - 'null'
          default: false
        claimName:
          description: Name for the PVC. Allow referencing it across workloads. If not provided, a name based on the workload name and scope will be auto-generated.
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: my-claim
          pattern: .*
        readOnly:
          description: Permit only read access to PVC.
          type:
          - boolean
          - 'null'
          default: false
        ephemeral:
          description: Use `true` to set PVC to ephemeral. If set to `true`, the PVC will be deleted when the workload is stopped. Not supported for inference workloads.
          type:
          - boolean
          - 'null'
          default: false
          example: false
        claimInfo:
          $ref: '#/components/schemas/ClaimInfo'
        dataSharing:
          description: use `true` to share the PVC data to all projects under the selected scope.
          type:
          - boolean
          - 'null'
          default: false
          example: false
      type:
      - object
      - 'null'
    DistributedInferenceSpecSpec:
      allOf:
      - $ref: '#/components/schemas/DistributedInferenceCommonSpec'
      - $ref: '#/components/schemas/DistributedInferenceLeaderSpecFields'
      - $ref: '#/components/schemas/DistributedInferenceWorkerSpecFields'
    Preemptibility:
      description: Specifies whether the workload can be preempted by higher-priority workloads. Valid values are preemptible and non-preemptible. If explicitly set, this value takes precedence. If not set, the system derives the preemptibility from the priorityClassName field, ensuring backward compatibility. Each workload type has a default preemptibility. To view the default preemptibility for each workload type, use the GET /workload-types endpoint.
      type:
      - string
      - 'null'
      minLength: 1
      enum:
      - preemptible
      - non-preemptible
    PvcAddedAttrValues:
      description: an optional array of key-values pairs that are written as annotations on the created PVC. the allowed attributes are determined according to the storage class configuration (see k8s-objects-tracker for further info).
      type: array
      items:
        $ref: '#/components/schemas/PvcAddedAttrValue'
    RunAsUid:
      description: The user id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset `runAsUid` field (optional). Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled.
      type:
      - integer
      - 'null'
      format: int64
      example: 500
    StorageInstanceName:
      properties:
        name:
          description: unique name to identify the instance. primarily used for policy locked rules.
          type:
          - string
          - 'null'
          minLength: 1
          example: storage-instance-a
      type:
      - object
      - 'null'
    WorkloadCreationMeta:
      required:
      - name
      - projectId
      - clusterId
      properties:
        name:
          $ref: '#/components/schemas/WorkloadName'
        useGivenNameAsPrefix:
          description: When true, the requested name will be treated as a prefix. The final name of the workload will be composed of the name followed by a random set of characters.
          type: boolean
          example: true
          default: false
        projectId:
          $ref: '#/components/schemas/ProjectId'
        clusterId:
          $ref: '#/components/schemas/ClusterId'
    WorkloadName:
      description: The name of the workload.
      type: string
      minLength: 1
      example: my-workload-name
      pattern: .*
    Image:
      description: Docker image name. For more information, see [Images](https://kubernetes.io/docs/concepts/containers/images). The image name is mandatory for creating a workload.
      type:
      - string
      - 'null'
      minLength: 1
      example: python:3.8
      pattern: .*
    SeccompProfileType:
      description: Indicates which kind of seccomp profile will be applied to the container. The options are a. `RuntimeDefault` - the container runtime default profile should be used. b. `Unconfined` - no profile should be applied. c. `Localhost` is not yet supported by Run:ai.
      type:
      - string
      - 'null'
      enum:
      - RuntimeDefault
      - Unconfined
      - Localhost
    EnvironmentVariable:
      description: Details of an environment variable which is populated into the container.
      properties:
        name:
          description: The name of the environment variable. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: HOME
          pattern: .*
        value:
          description: The value of the environment variable. (mutually exclusive with secret, userCredential, configMap and podFieldRef)
          type:
          - string
          - 'null'
          example: /home/my-folder
          pattern: .*
        secret:
          $ref: '#/components/schemas/EnvironmentVariableSecret'
        configMap:
          $ref: '#/components/schemas/EnvironmentVariableConfigMap'
        podFieldRef:
          $ref: '#/components/schemas/EnvironmentVariablePodFieldReference'
        userCredential:
          $ref: '#/components/schemas/EnvironmentVariableUserCredential'
        exclude:
          description: Use 'true' in case the environment variable is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
        description:
          description: Description of the environment variable.
          type:
          - string
          - 'null'
          example: Home directory of the user.
          pattern: .*
      type:
      - object
      - 'null'
    DistributedInferenceStartupPolicy:
      description: "Determines when the worker pods should start during workload initialization. \n - `LeaderCreated`: Workers start after the leader pod is created.\n - `LeaderReady`: Workers start only after the leader pod is ready.\n"
      type:
      - string
      - 'null'
      enum:
      - LeaderCreated
      - LeaderReady
      default: LeaderCreated
    PvcAccessModes:
      description: Default access mode(s) applied to newly created PVCs unless explicitly overridden.
      properties:
        readWriteOnce:
          description: Mount the volume as read/write by a single node.
          type:
          - boolean
          - 'null'
          default: true
        readOnlyMany:
          description: Mount the volume as read-only by many nodes.
          type:
          - boolean
          - 'null'
          default: false
        readWriteMany:
          description: Mount the volume as read/write by many nodes.
          type:
          - boolean
          - 'null'
          default: false
      type:
      - object
      - 'null'
    EnvironmentVariables:
      description: Set of environment variables to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/EnvironmentVariable'
    Probe:
      type:
      - object
      - 'null'
      properties:
        initialDelaySeconds:
          description: Number of seconds after the container has started before liveness or readiness probes are initiated.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 0
        periodSeconds:
          description: How often (in seconds) to perform the probe.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        timeoutSeconds:
          description: Number of seconds after which the probe times out.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        successThreshold:
          description: Minimum consecutive successes for the probe to be considered successful after having failed.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        failureThreshold:
          description: When a probe fails, the number of times to try before giving up.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        handler:
          $ref: '#/components/schemas/ProbeHandler'
    PodAffinity:
      description: Pod affinity scheduling rules (e.g. co-locate this workload in the same node, zone, etc. as some other workloads).
      type:
      - object
      - 'null'
      properties:
        type:
          $ref: '#/components/schemas/PodAffinityType'
        key:
          description: The label key to use. (mandatory)
          type:
          - string
          - 'null'
          pattern: .*
    CreateHomeDir:
      description: When set to `true`, creates a home directory for the container.
      type:
      - boolean
      - 'null'
    UpdateSpec:
      description: The specifications of the inference to be updated.
      properties:
        spec:
          allOf:
          - type:
            - object
            - 'null'
            properties:
              replicas:
                description: Specifies the number of leader-worker sets to deploy. Each replica represents a group consisting of one leader pod and multiple worker pods. For example, setting replicas to 3 will create 3 independent groups, each with its own leader and corresponding set of workers.
                type:
                - integer
                - 'null'
                format: int32
                minimum: 0
                maximum: 1000
                example: 2
    ProjectId:
      description: The id of the project.
      type: string
      example: 1
      pattern: .*
    DistributedInferenceReplicasField:
      properties:
        replicas:
          default: 1
          description: "Specifies the number of leader-worker sets to deploy. Each replica represents a group consisting of one leader pod and multiple worker pods. \nFor example, setting replicas: 3 will create 3 independent groups, each with its own leader and corresponding set of workers.\n"
          type:
          - integer
          - 'null'
          format: int32
          minimum: 0
          maximum: 1000
          example: 2
    Tolerations:
      description: Set of tolerations to apply to the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Toleration'
    DistributedInference:
      allOf:
      - $ref: '#/components/schemas/WorkloadMeta1'
      - $ref: '#/components/schemas/DistributedInferenceSpec'
    ExtendedResource:
      description: Quantity of an extended resource.
      properties:
        resource:
          description: The name of the extended resource (mandatory)
          type:
          - string
          - 'null'
          example: hardware-vendor.example/foo
          minLength: 1
          pattern: .*
        quantity:
          description: The requested quantity for the resource.
          type:
          - string
          - 'null'
          example: 2
          minLength: 1
          pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
        exclude:
          description: Use 'true' in case the extended resource is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
      type:
      - object
      - 'null'
    DistributedInferenceServingPort:
      description: Defines the configuration for the inference serving endpoint. This determines how applications or services can send inference requests to the workload.
      allOf:
      - $ref: '#/components/schemas/DistributedInferenceServingPortContainerAndProtocol'
      - $ref: '#/components/schemas/DistributedInferenceServingPortAccess'
      type:
      - object
      - 'null'
    DistributedInferenceServingPortProtocol:
      description: The protocol used to access the port.
      type:
      - string
      - 'null'
      enum:
      - http
      default: http
    PvcItems:
      description: Set of pvc persistent volume claims to use in the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/PvcInstance'
    DistributedInferenceLeaderWorkerSpec1:
      properties:
        annotations:
          $ref: '#/components/schemas/Annotations'
        args:
          $ref: '#/components/schemas/Args'
        command:
          $ref: '#/components/schemas/Command'
        compute:
          properties:
            cpuCoreLimit:
              $ref: '#/components/schemas/CpuCoreLimit'
            cpuCoreRequest:
              $ref: '#/components/schemas/CpuCoreRequest'
            cpuMemoryLimit:
              $ref: '#/components/schemas/CpuMemoryLimit'
            cpuMemoryRequest:
              $ref: '#/components/schemas/CpuMemoryRequest'
            extendedResources:
              $ref: '#/components/schemas/ExtendedResources'
            gpuDevicesRequest:
              $ref: '#/components/schemas/GpuDevicesRequest'
            gpuMemoryLimit:
              $ref: '#/components/schemas/GpuMemoryLimit'
            gpuMemoryRequest:
              $ref: '#/components/schemas/GpuMemoryRequest'
            gpuPortionLimit:
              $ref: '#/components/schemas/GpuPortionLimit'
            gpuPortionRequest:
              $ref: '#/components/schemas/GpuPortionRequest'
            gpuRequestType:
              $ref: '#/components/schemas/GpuRequestType'
            largeShmRequest:
              $ref: '#/components/schemas/LargeShmRequest'
          type:
          - object
          - 'null'
        createHomeDir:
          $ref: '#/components/schemas/CreateHomeDir'
        environmentVariables:
          $ref: '#/components/schemas/EnvironmentVariables'
        image:
          $ref: '#/components/schemas/Image'
        imagePullPolicy:
          $ref: '#/components/schemas/ImagePullPolicy'
        imagePullSecrets:
          $ref: '#/components/schemas/ImagePullSecrets'
        labels:
          $ref: '#/components/schemas/Labels'
        nodeAffinityRequired:
          $ref: '#/components/schemas/NodeAffinityRequired'
        nodeType:
          $ref: '#/components/schemas/NodeType3'
        podAffinity:
          $ref: '#/components/schemas/PodAffinity'
        probes:
          $ref: '#/components/schemas/Probes'
        security:
          properties:
            capabilities:
              $ref: '#/components/schemas/Capabilities'
            readOnlyRootFilesystem:
              $ref: '#/components/schemas/ReadOnlyRootFileSystem'
            runAsGid:
              $ref: '#/components/schemas/RunAsGid'
            runAsNonRoot:
              $ref: '#/components/schemas/RunAsNonRoot'
            runAsUid:
              $ref: '#/components/schemas/RunAsUid'
            seccompProfileType:
              $ref: '#/components/schemas/SeccompProfileType'
            supplementalGroups:
              $ref: '#/components/schemas/SupplementalGroups'
            uidGidSource:
              $ref: '#/components/schemas/UidGidSource'
          type:
          - object
          - 'null'
        storage:
          properties:
            configMapVolume:
              $ref: '#/components/schemas/ConfigMapItems'
            emptyDirVolume:
              $ref: '#/components/schemas/EmptyDirItems'
            pvc:
              $ref: '#/components/schemas/PvcItems'
            secretVolume:
              $ref: '#/components/schemas/SecretItems1'
          type:
          - object
          - 'null'
        tolerations:
          $ref: '#/components/schemas/Tolerations'
        workingDir:
          $ref: '#/components/schemas/WorkingDir'
      type: object
    Args:
      description: Arguments to the command that the container running the workload executes.
      type:
      - string
      - 'null'
      minLength: 1
      example: -x my-script.py
      pattern: .*
    DistributedInferenceServingPortAccess:
      properties:
        authorizationType:
          $ref: '#/components/schemas/DistributedInferenceServingPortAccessAuthorizationTypeEnum'
        authorizedUsers:
          description: A list of users and service accounts allowed to send inference requests to the serving endpoint. `Note:` Cannot be used together with authorizedGroups.
          type:
          - array
          - 'null'
          items:
            type: string
            pattern: .*
          example:
          - user.a@example.com
          - user.b@example.com
        authorizedGroups:
          description: A list of user groups allowed to send inference requests to the serving endpoint. `Note:` Cannot be used together with authorizedUsers.
          type:
          - array
          - 'null'
          items:
            type: string
            pattern: .*
          example:
          - group-a
          - group-b
        exposeExternally:
          description: Indicates whether the inference serving endpoint should be accessible outside the cluster. If set to true, the endpoint will be exposed externally. To enable external access, your administrator must configure the cluster as described in the [inference requirements](https://run-ai-docs.nvidia.com/saas/getting-started/installation/system-requirements#inference). section.
          type:
          - boolean
          - 'null'
          default: true
        exposedUrl:
          description: The custom URL to use for the serving port. If empty (default), an autogenerated URL will be used.
          type:
          - string
          - 'null'
          pattern: .*
    EnvironmentVariableSecret:
      description: Details of the secret and key use to populate the environment variable
      properties:
        name:
          description: The name of the secret resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: postgress_secret
          pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
        key:
          description: The key in the secret resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: POSTGRES_PASSWORD
          pattern: .*
      type:
      - object
      - 'null'
    Command:
      description: A command to the server as the entry point of the container running the workload.
      type:
      - string
      - 'null'
      minLength: 1
      example: python
      pattern: .

# --- truncated at 32 KB (61 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/runai/refs/heads/main/openapi/runai-distributed-inferences-api-openapi.yml