NVIDIA Run:ai Inferences API

Inference workloads deploy trained models into a production environment to generate predictions from live data. These workloads are prioritized over Trainings and Workspaces during scheduling. NVIDIA Run:ai Inference workloads support auto-scaling to maintain service-level agreements (SLAs) by dynamically adjusting resources as demand changes.

OpenAPI Specification

runai-inferences-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  version: latest
  description: '# Introduction


    The NVIDIA Run:ai Control-Plane API reference is a guide that provides an easy-to-use programming interface for adding various tasks to your application, including workload submission, resource management, and administrative operations.


    NVIDIA Run:ai APIs are accessed using *bearer tokens*. To obtain a token, you need to create a **Service account** through the NVIDIA Run:ai user interface.

    To create a service account, in your UI, go to Access → Service Accounts (for organization-level service accounts) or User settings → Access Keys (for user access keys), and create a new one.


    After you have created a new service account, you will need to assign it access rules.

    To assign access rules to the service account, see [Create access rules](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/accessrules#create-or-delete-rules).

    Make sure you assign the correct rules to your service account. Use the [Roles](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/roles) to assign the correct access rules.


    To get your access token, follow the instructions in [Request a token](https://run-ai-docs.nvidia.com/saas/reference/api/rest-auth/#request-an-api-token).

    '
  title: NVIDIA Run:ai Access Keys Inferences API
  x-logo:
    url: https://api.redocly.com/registry/raw/runai-xq8/saas/latest/public/runai-logo-api.png
    altText: NVIDIA Run:ai
    href: https://run.ai
  license:
    name: NVIDIA Run:ai
    url: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-license-agreement/
servers:
- url: https://app.run.ai
security:
- bearerAuth: []
tags:
- name: Inferences
  description: Inference workloads deploy trained models into a production environment to generate predictions from live data. These workloads are prioritized over Trainings and Workspaces during scheduling. NVIDIA Run:ai Inference workloads support auto-scaling to maintain service-level agreements (SLAs) by dynamically adjusting resources as demand changes.
paths:
  /api/v1/workloads/inferences:
    post:
      summary: Create an inference.
      operationId: create_inference1
      description: Create an inference using container related fields.
      tags:
      - Inferences
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/InferenceCreationRequest'
      responses:
        '202':
          description: Request completed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Inference1'
        '400':
          $ref: '#/components/responses/400BadRequest'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
  /api/v1/workloads/inferences/{workloadId}:
    delete:
      summary: Delete an inference.
      operationId: delete_inference
      description: Delete an inference using a workload id.
      tags:
      - Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      responses:
        '202':
          $ref: '#/components/responses/202Accepted'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
    get:
      summary: Get inference data.
      operationId: get_inference
      description: Retrieve inference details using a workload id.
      tags:
      - Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      responses:
        '200':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Inference1'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
    patch:
      summary: Update inference spec. [Experimental]
      operationId: update_inference_spec
      description: Update the specification of an existing inference workload.
      tags:
      - Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/InferenceUpdateRequest'
      responses:
        '202':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Inference1'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
  /api/v1/workloads/inferences/{workloadId}/metrics:
    get:
      summary: Get inference metrics data.
      description: Retrieve inference metrics data by id. Supported from control-plane version 2.18 or later.
      operationId: get_inference_workload_metrics
      tags:
      - Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      - $ref: '#/components/parameters/InferenceWorkloadMetricTypes'
      - $ref: '#/components/parameters/StartRequired'
      - $ref: '#/components/parameters/EndRequired'
      - $ref: '#/components/parameters/NumberOfSamples'
      responses:
        '200':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/MetricsResponse'
            text/csv: {}
        '207':
          description: Partial success.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/MetricsResponse'
        '400':
          $ref: '#/components/responses/400BadRequest'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
  /api/v1/workloads/inferences/{workloadId}/pods/{podId}/metrics:
    get:
      summary: Get inference pod's metrics data.
      description: Retrieve inference metrics pod's data by workload and pod id. Supported from control-plane version 2.18 or later.
      operationId: get_inference_workload_pod_metrics
      tags:
      - Inferences
      parameters:
      - $ref: '#/components/parameters/WorkloadId'
      - $ref: '#/components/parameters/PodId'
      - $ref: '#/components/parameters/InferencePodMetricTypes'
      - $ref: '#/components/parameters/StartRequired'
      - $ref: '#/components/parameters/EndRequired'
      - $ref: '#/components/parameters/NumberOfSamples'
      responses:
        '200':
          description: Executed successfully.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/MetricsResponse'
            text/csv: {}
        '207':
          description: Partial success.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/MetricsResponse'
        '400':
          $ref: '#/components/responses/400BadRequest'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
components:
  schemas:
    MetricThresholdField:
      properties:
        metricThreshold:
          description: The threshold to use with the specified metric for autoscaling. Mandatory if metric is specified
          type:
          - integer
          - 'null'
          format: int32
    PodAffinity:
      description: Pod affinity scheduling rules (e.g. co-locate this workload in the same node, zone, etc. as some other workloads).
      type:
      - object
      - 'null'
      properties:
        type:
          $ref: '#/components/schemas/PodAffinityType'
        key:
          description: The label key to use. (mandatory)
          type:
          - string
          - 'null'
          pattern: .*
    InitialReplicasField:
      properties:
        initialReplicas:
          description: The number of replicas to run when initializing the workload for the first time. Defaults to minReplicas, or to 1 if minReplicas is set to 0
          type:
          - integer
          - 'null'
          format: int32
          minimum: 0
    MetricsResponse:
      type: object
      required:
      - measurements
      properties:
        measurements:
          type: array
          items:
            $ref: '#/components/schemas/MeasurementResponse'
    NodeAffinityRequired:
      type:
      - object
      - 'null'
      description: If the affinity requirements specified by this field are not met at scheduling time, the pod will not be scheduled onto the node. If the affinity requirements specified by this field cease to be met at some point during pod execution (e.g. due to an update), the system may or may not try to eventually evict the pod from its node.
      properties:
        nodeSelectorTerms:
          description: A list of node selector terms. The terms are ORed.
          type: array
          items:
            $ref: '#/components/schemas/NodeSelectorTerm'
    Capability:
      type: string
      enum:
      - AUDIT_CONTROL
      - AUDIT_READ
      - AUDIT_WRITE
      - BLOCK_SUSPEND
      - CHOWN
      - DAC_OVERRIDE
      - DAC_READ_SEARCH
      - FOWNER
      - FSETID
      - IPC_LOCK
      - IPC_OWNER
      - KILL
      - LEASE
      - LINUX_IMMUTABLE
      - MAC_ADMIN
      - MAC_OVERRIDE
      - MKNOD
      - NET_ADMIN
      - NET_BIND_SERVICE
      - NET_BROADCAST
      - NET_RAW
      - SETGID
      - SETFCAP
      - SETPCAP
      - SETUID
      - SYS_ADMIN
      - SYS_BOOT
      - SYS_CHROOT
      - SYS_MODULE
      - SYS_NICE
      - SYS_PACCT
      - SYS_PTRACE
      - SYS_RAWIO
      - SYS_RESOURCE
      - SYS_TIME
      - SYS_TTY_CONFIG
      - SYSLOG
      - WAKE_ALARM
    SupplementalGroups:
      description: Comma separated list of groups that the user running the container belongs to, in addition to the group indicated by runAsGid. Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled. Using an empty string implies reverting the supplementary groups of the image.
      type:
      - string
      - 'null'
      example: 2,3,5,8
      pattern: .*
    InitializationTimeoutField:
      properties:
        initializationTimeoutSeconds:
          description: Use `servingConfiguration.initializationTimeoutSeconds` instead.  If this field is set, it will be ignored and the value under `servingConfiguration` will be used. The maximum amount of time (in seconds) to wait for the container to become ready.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
          deprecated: true
    Tolerations:
      description: Set of tolerations to apply to the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Toleration'
    ImagePullPolicy:
      description: Image pull policy. Defaults to `Always` if `:latest` tag is specified, otherwise it is `IfNotPresent`.
      type:
      - string
      - 'null'
      minLength: 1
      enum:
      - Always
      - Never
      - IfNotPresent
    UidGidSource:
      description: Indicate the way to determine the user and group ids of the container. The options are a. `fromTheImage` - user and group ids are determined by the docker image that the container runs. this is the default option. b. `custom` - user and group ids can be specified in the environment asset and/or the workload creation request. c. `idpToken` - user and group IDs are automatically taken from the identity provider (IdP) token (available only in SSO-enabled installations). For more information, see [User Identity](https://docs.run.ai/latest/admin/runai-setup/config/non-root-containers/).
      type:
      - string
      - 'null'
      enum:
      - fromTheImage
      - fromIdpToken
      - custom
    Label:
      description: Label details to be populated into the container.
      properties:
        name:
          description: The name of the label (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: stage
          pattern: .*
        value:
          description: The value of the label.
          type:
          - string
          - 'null'
          example: initial-research
          pattern: .*
        exclude:
          description: Use 'true' in case the label is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
      type:
      - object
      - 'null'
    ProbeHandler:
      description: The action taken to determine the health of the container. (mandatory)
      type:
      - object
      - 'null'
      properties:
        httpGet:
          description: An action based on HTTP Get requests.
          type: object
          properties:
            path:
              description: Path to access on the HTTP server, defaults to /.
              type:
              - string
              - 'null'
              pattern: ^(\x2F[a-zA-Z0-9\-_.\x2F]*)?$
              example: /
            port:
              description: Number of the port to access on the container.
              type:
              - integer
              - 'null'
              format: int32
              minimum: 1
              maximum: 65535
            host:
              description: Host name to connect to, defaults to the pod IP.
              type:
              - string
              - 'null'
              format: hostname
              example: example.com
              pattern: .*
            scheme:
              $ref: '#/components/schemas/ProbeHandlerScheme'
    EnvironmentVariablePodFieldReference:
      description: Details of the field-reference and key use to populate the environment variable
      properties:
        path:
          description: The field path resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: metadata.name
          pattern: .*
      type:
      - object
      - 'null'
    ImagePullSecrets:
      description: A list of references to Kubernetes secrets in the same namespace used for pulling container images.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/ImagePullSecret'
    Labels:
      description: Set of labels to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Label'
    CpuMemoryLimit:
      description: Limitations on the CPU memory to allocate for this workload (1G, 20M, .etc). The system guarantees that this workload will not be able to consume more than this amount of memory. The workload will receive an error when trying to allocate more memory than this limit.
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 30M
    SecretFieldsNonUpdatable:
      properties:
        secret:
          description: The name of the Secret resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
      type:
      - object
      - 'null'
    ServingPortAccessAuthorizationTypeEnum:
      type:
      - string
      - 'null'
      enum:
      - public
      - authenticatedUsers
      - authorizedUsers
      - authorizedGroups
      - authorizedUsersOrGroups
      description: 'Specifies who can send inference requests to the serving endpoint:


        Possible values:

        - `public`: No authorization is required. (Default)

        - `authenticatedUsers`: Any NVIDIA Run:ai authenticated user and service account can send requests.

        - `authorizedUsers`: Only users listed in the authorizedUsers field can send requests.

        - `authorizedGroups`: Only members of user groups listed in the authorizedGroups field can send requests.

        - `authorizedUsersOrGroups`: Requires either authorizedUsers or authorizedGroups to be provided; if neither is set, or if both are set, a mutual exclusion error is reported. Supported from cluster version 2.19.

        '
    GitInstance:
      allOf:
      - $ref: '#/components/schemas/StorageInstanceName'
      - $ref: '#/components/schemas/GitCommon'
      - $ref: '#/components/schemas/GitPassword'
      - $ref: '#/components/schemas/ExcludeField'
      - type:
        - object
        - 'null'
        properties:
          secretRef:
            $ref: '#/components/schemas/GitSecretRef'
      type:
      - object
      - 'null'
    InferenceUpdateSpec:
      description: The specifications of the inference to be updated.
      properties:
        spec:
          allOf:
          - properties:
              args:
                $ref: '#/components/schemas/Args'
              category:
                $ref: '#/components/schemas/Category'
              command:
                $ref: '#/components/schemas/Command'
              compute:
                properties:
                  cpuCoreLimit:
                    $ref: '#/components/schemas/CpuCoreLimit'
                  cpuCoreRequest:
                    $ref: '#/components/schemas/CpuCoreRequest'
                  cpuMemoryLimit:
                    $ref: '#/components/schemas/CpuMemoryLimit'
                  cpuMemoryRequest:
                    $ref: '#/components/schemas/CpuMemoryRequest'
                  extendedResources:
                    $ref: '#/components/schemas/ExtendedResources'
                  gpuDevicesRequest:
                    $ref: '#/components/schemas/GpuDevicesRequest'
                  gpuMemoryLimit:
                    $ref: '#/components/schemas/GpuMemoryLimit'
                  gpuMemoryRequest:
                    $ref: '#/components/schemas/GpuMemoryRequest'
                  gpuPortionLimit:
                    $ref: '#/components/schemas/GpuPortionLimit'
                  gpuPortionRequest:
                    $ref: '#/components/schemas/GpuPortionRequest'
                  gpuRequestType:
                    $ref: '#/components/schemas/GpuRequestType'
                  largeShmRequest:
                    $ref: '#/components/schemas/LargeShmRequest'
                type:
                - object
                - 'null'
              createHomeDir:
                $ref: '#/components/schemas/CreateHomeDir'
              environmentVariables:
                $ref: '#/components/schemas/EnvironmentVariables'
              image:
                $ref: '#/components/schemas/Image'
              imagePullPolicy:
                $ref: '#/components/schemas/ImagePullPolicy'
              imagePullSecrets:
                $ref: '#/components/schemas/ImagePullSecrets'
              nodeAffinityRequired:
                $ref: '#/components/schemas/NodeAffinityRequired'
              nodePools:
                $ref: '#/components/schemas/NodePools'
              nodeType:
                $ref: '#/components/schemas/NodeType3'
              podAffinity:
                $ref: '#/components/schemas/PodAffinity'
              preemptibility:
                $ref: '#/components/schemas/Preemptibility'
              priorityClass:
                $ref: '#/components/schemas/PriorityClass'
              probes:
                $ref: '#/components/schemas/Probes'
              workingDir:
                $ref: '#/components/schemas/WorkingDir'
            type:
            - object
            - 'null'
          - $ref: '#/components/schemas/InferenceUpdateSpecAutoscaling'
          - $ref: '#/components/schemas/InferenceUpdateSpecServingConfiguration'
    NfsInstance:
      allOf:
      - $ref: '#/components/schemas/StorageInstanceName'
      - $ref: '#/components/schemas/Nfs'
      - $ref: '#/components/schemas/ExcludeField'
      type:
      - object
      - 'null'
    PortServiceType:
      description: The service type of the port (mandatory).
      type:
      - string
      - 'null'
      enum:
      - LoadBalancer
      - NodePort
      - ClusterIP
    EnvironmentVariableUserCredential:
      description: Defines a reference to a user-created credential and a specific key within that credential whose value will populate the environment variable. User credentials can only be accessed by the user who created them.
      properties:
        name:
          description: The name of the user credential.  (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: my_postgres_user_and_password
        key:
          description: The key in the user credential resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: POSTGRES_PASSWORD
      type:
      - object
      - 'null'
    Args:
      description: Arguments to the command that the container running the workload executes.
      type:
      - string
      - 'null'
      minLength: 1
      example: -x my-script.py
      pattern: .*
    Phase:
      type: string
      enum:
      - Creating
      - Initializing
      - Resuming
      - Pending
      - Deleting
      - Running
      - Updating
      - Stopped
      - Stopping
      - Degraded
      - Failed
      - Completed
      - Terminating
      - Unknown
    SeccompProfileType:
      description: Indicates which kind of seccomp profile will be applied to the container. The options are a. `RuntimeDefault` - the container runtime default profile should be used. b. `Unconfined` - no profile should be applied. c. `Localhost` is not yet supported by Run:ai.
      type:
      - string
      - 'null'
      enum:
      - RuntimeDefault
      - Unconfined
      - Localhost
    AutoScalingCommonFields:
      allOf:
      - $ref: '#/components/schemas/MetricThresholdPercentageField'
      - $ref: '#/components/schemas/InferencesMinReplicasField'
      - $ref: '#/components/schemas/InferencesMaxReplicasField'
      - $ref: '#/components/schemas/InitialReplicasField'
      - $ref: '#/components/schemas/ActivationReplicasField'
      - $ref: '#/components/schemas/ConcurrencyHardLimitField'
      - $ref: '#/components/schemas/ScaleToZeroRetentionField'
      - $ref: '#/components/schemas/ScaleDownDelayField'
      - $ref: '#/components/schemas/InitializationTimeoutField'
      type:
      - object
      - 'null'
    EmptyDir:
      properties:
        path:
          description: Local path within the workload to which the EmptyDir volume will be mapped. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: /mnt/emptydir
          pattern: .*
        medium:
          description: The type of storage medium for the volume. Use "Memory" for memory-backed storage, or leave empty for disk-backed storage.
          type:
          - string
          - 'null'
          minLength: 1
          pattern: .*
        sizeLimit:
          description: The total amount of local storage or memory required for the emptyDir volume. Specify using Kubernetes quantity format (e.g., 1G, 500Mi).
          type:
          - string
          - 'null'
          pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
          example: 1G
      type:
      - object
      - 'null'
    PodAffinityType:
      description: The affinity type, required or preferred. (mandatory)
      type:
      - string
      - 'null'
      enum:
      - Required
      - Preferred
    Annotations:
      description: Set of annotations to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Annotation'
    HostPathItems:
      description: Set of host paths to use in the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/HostPathInstance'
    GitPassword:
      properties:
        passwordSecret:
          description: Secret containing the credentials of the repository (needed for non public repository which requires authentication). (deprecated)
          type:
          - string
          - 'null'
          minLength: 1
          example: my-password-secret
        secretKeyOfUser:
          description: The key to use for loading the user name from the secret. The default is `User`. (deprecated)
          type:
          - string
          - 'null'
          minLength: 1
          example: User
        secretKeyOfPassword:
          description: The key to use for loading the password from the secret. The default is `Password`. (deprecated)
          type:
          - string
          - 'null'
          minLength: 1
          example: Password
      type:
      - object
      - 'null'
    Annotation:
      description: Annotation details to be populated into the container.
      properties:
        name:
          description: The name of the annotation (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: billing
          pattern: .*
        value:
          description: The value of the annotation.
          type:
          - string
          - 'null'
          example: my-billing-unit
          pattern: .*
        exclude:
          description: Use 'true' in case the annotation is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          default: false
          example: false
      type:
      - object
      - 'null'
    DataVolumeInstance:
      allOf:
      - $ref: '#/components/schemas/DataVolume'
      - $ref: '#/components/schemas/ExcludeField'
      type:
      - object
      - 'null'
    RunAsUid:
      description: The user id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset `runAsUid` field (optional). Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled.
      type:
      - integer
      - 'null'
      format: int64
      example: 500
    ReadOnlyRootFileSystem:
      description: If true, mounts the container's root filesystem as read-only.
      type:
      - boolean
      - 'null'
      example: false
    MeasurementResponse:
      type: object
      required:
      - type
      - values
      properties:
        type:
          type: string
          description: specifies what data returned
          example: ALLOCATED_GPU
        labels:
          type:
          - object
          - 'null'
          description: labels of the metric measurement
          example: '{''gpu'': ''3''}'
          additionalProperties:
            type: string
        values:
          type:
          - array
          - 'null'
          items:
            type: object
            required:
            - value
            - timestamp
            properties:
              value:
                type: string
                example: '85'
              timestamp:
                type:
                - string
                - 'null'
                format: date-time
                example: '2023-06-06 12:09:18.211'
    PvcVolumeMode:
      description: Default volume mode for the PVC. Choose between Filesystem (default) or Block.
      type:
      - string
      - 'null'
      enum:
      - Filesystem
      - Block
    ConfigMapInstance:
      allOf:
      - $ref: '#/components/schemas/StorageInstanceName'
      - $ref: '#/components/schemas/ConfigMap'
      - $ref: '#/components/schemas/ExcludeField'
      type:
      - object
      - 'null'
    RelatedUrl:
      description: A URL that is related to the workload. For example, a URL to an external server providing statistics or logging about the workload.
      properties:
        url:
          description: The URL for connecting an external service related to the workload. (mandatory)
          type:
          - string
          - 'null'
          example: https://my-url.com
          pattern: .*
        type:
          description: The type of service that the url provides. For example, wandb (Weights & Biases). (mandatory)
          type:
          - string
          - 'null'
          example: wandb
          pattern: .*
        name:
          description: Unique name to identify the instance. primarily used for policy locked rules.
          type:
          - string
          - 'null'
          example: url-instance-a
          pattern: .*
        exclude:
          description: Use 'true' in case the item is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          default: false
          example: false
      type:
      - object
      - 'null'
    NodePools:
      description: A prioritized list of node pools for the scheduler to run the workload on. The scheduler will always try to use the first node pool before moving to the next one if the first is not available.
      type:
      - array
      - 'null'
      items:
        type: string
        pattern: .*
      example:
      - my-node-pool-a
      - my-node-pool-b
    PvcClaimSize:
      description: 'Requested size for the PVC. Mandatory when existingPvc is false. Recommended sizes: TB/GB/MB/TIB/GIB/MIB'
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 1G
    ExposedUrl:
      description: A URL for accessing the workload.
      properties:
        container:
          description: The port that the container running the workload exposes. (mandatory)
          type:
          - integer
          - 'null'
          format: int32
          example: 8080
        url:
          description: The URL for connecting to the container port. If not specified, the URL will be auto-generated by the system..
          type:
          - string
          - 'null'
          pattern: .*
          example: https://my-url.com
        authorizationType:
          $ref: '#/components/schemas/AuthorizationType'
        authorizedUsers:
          description: List of users or service accounts that are allowed to access the URL. Note that authorizedUsers and authorizedGroups are mutually exclusive.
          type:
          - array
          - 'null'
          items:
            type: string
            pattern: .*
          example:
          - user-a
          - user-b
        authorizedGroups:
          description: List of groups that are allowed to access the URL. Note that authorizedUsers and authorizedGroups are mutually exclusive.
          type:
          - array
          - 'null'
          items:
            type: string
            pattern: .*
          example:
          - group-a
          - group-b
        toolType:
          description: The tool type that runs on this container port.
          type:
          - string
          - 'null'
          example: jupyter
          pattern: .*
        toolName:
          description: A name describing the tool that runs on this url.
          type:
          - string
          - 'null'
          example: my-pytorch
          pattern: .*
        name:
          description: Unique name to identify the 

# --- truncated at 32 KB (86 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/runai/refs/heads/main/openapi/runai-inferences-api-openapi.yml