NVIDIA Run:ai NVIDIA NIM API

The NVIDIA NIM API provides endpoints to create and manage workloads that deploy NVIDIA Inference Microservices (NIM) through the NIM Operator. These workloads package optimized NVIDIA model servers and run as managed services on the NVIDIA Run:ai platform. Each request includes NVIDIA Run:ai scheduling metadata (for example, project, priority, and category) and a NIM service specification that defines the container image, compute resources, environment variables, storage, and networking configuration. Once submitted, NVIDIA Run:ai handles scheduling, orchestration, and lifecycle management of the NIM service to ensure reliable and efficient model serving.

Operations 3

POST /api/v2/workloads/nim-services Create a NVIDIA NIM service. [Experimental] #
GET /api/v2/workloads/nim-services/{WorkloadV2Id} Get a NVIDIA NIM service. [Experimental] #
PATCH /api/v2/workloads/nim-services/{WorkloadV2Id} Update NVIDIA NIM service spec. [Experimental] #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/runai-nvidia-nim-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

runai-nvidia-nim-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  version: latest
  description: '# Introduction


    The NVIDIA Run:ai Control-Plane API reference is a guide that provides an easy-to-use programming interface for adding various tasks to your application, including workload submission, resource management, and administrative operations.


    NVIDIA Run:ai APIs are accessed using *bearer tokens*. To obtain a token, you need to create a **Service account** through the NVIDIA Run:ai user interface.

    To create a service account, in your UI, go to Access → Service Accounts (for organization-level service accounts) or User settings → Access Keys (for user access keys), and create a new one.


    After you have created a new service account, you will need to assign it access rules.

    To assign access rules to the service account, see [Create access rules](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/accessrules#create-or-delete-rules).

    Make sure you assign the correct rules to your service account. Use the [Roles](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/roles) to assign the correct access rules.


    To get your access token, follow the instructions in [Request a token](https://run-ai-docs.nvidia.com/saas/reference/api/rest-auth/#request-an-api-token).

    '
  title: NVIDIA Run:ai Access Keys NVIDIA NIM API
  x-logo:
    url: https://api.redocly.com/registry/raw/runai-xq8/saas/latest/public/runai-logo-api.png
    altText: NVIDIA Run:ai
    href: https://run.ai
  license:
    name: NVIDIA Run:ai
    url: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-license-agreement/
servers:
- url: https://app.run.ai
security:
- bearerAuth: []
tags:
- name: NVIDIA NIM
  description: 'The NVIDIA NIM API provides endpoints to create and manage workloads that deploy NVIDIA Inference Microservices (NIM) through the NIM Operator. These workloads package optimized NVIDIA model servers and run as managed services on the NVIDIA Run:ai platform.

    Each request includes NVIDIA Run:ai scheduling metadata (for example, project, priority, and category) and a NIM service specification that defines the container image, compute resources, environment variables, storage, and networking configuration. Once submitted, NVIDIA Run:ai handles scheduling, orchestration, and lifecycle management of the NIM service to ensure reliable and efficient model serving.

    '
paths:
  /api/v2/workloads/nim-services:
    post:
      summary: Create a NVIDIA NIM service. [Experimental]
      description: Create a NVIDIA NIM service
      operationId: create_nim_service
      tags:
      - NVIDIA NIM
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/NimServiceCreateRequest'
      responses:
        '202':
          description: Workload creation accepted
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/NimServiceResponse'
        '400':
          $ref: '#/components/responses/400SubmissionErrorV2'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '409':
          $ref: '#/components/responses/409Conflict'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
  /api/v2/workloads/nim-services/{WorkloadV2Id}:
    get:
      summary: Get a NVIDIA NIM service. [Experimental]
      description: Retrieve details of a specific NVIDIA NIM service, by id
      operationId: get_nim_service_by_id
      tags:
      - NVIDIA NIM
      parameters:
      - $ref: '#/components/parameters/WorkloadV2Id'
      responses:
        '200':
          description: Successfully retrieved the workload
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/NimServiceResponse'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
    patch:
      summary: Update NVIDIA NIM service spec. [Experimental]
      operationId: update_nim_service_spec
      description: Update the specification of an existing NVIDIA NIM service.
      tags:
      - NVIDIA NIM
      parameters:
      - $ref: '#/components/parameters/WorkloadV2Id'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/NimServiceUpdateRequest'
      responses:
        '202':
          description: Workload update request accepted
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/NimServiceResponse'
        '401':
          $ref: '#/components/responses/401Unauthorized'
        '403':
          $ref: '#/components/responses/403Forbidden'
        '404':
          $ref: '#/components/responses/404NotFound'
        '500':
          $ref: '#/components/responses/500InternalServerError'
        '503':
          $ref: '#/components/responses/503ServiceUnavailable'
components:
  schemas:
    NimServicePvcFields:
      properties:
        existingPvc:
          description: Verify existing PVC. PVC is assumed to exist when set to `true`. If set to `false`, the PVC will be created, if it does not exist.
          type:
          - boolean
          - 'null'
          default: false
        claimName:
          description: Name for the PVC. Allow referencing it across workloads. If not provided, a name based on the workload name and scope will be auto-generated.
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: my-claim
          pattern: .*
        readOnly:
          description: Permit only read access to PVC.
          type:
          - boolean
          - 'null'
          default: false
        claimInfo:
          $ref: '#/components/schemas/ClaimInfo'
      type:
      - object
      - 'null'
    EnvironmentVariableConfigMap:
      description: Details of the configMap and key use to populate the environment variable
      properties:
        name:
          description: The name of the config-map resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: my-config-map
          pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
        key:
          description: The key in the config-map resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: MY_POSTGRES_SCHEMA
          pattern: .*
      type:
      - object
      - 'null'
    NimServiceServingPort:
      description: A port for accessing the inference service
      properties:
        serviceType:
          $ref: '#/components/schemas/ServingPortServiceType'
        port:
          $ref: '#/components/schemas/ServingPortPort'
        grpcPort:
          $ref: '#/components/schemas/ServingPortGrpcPort'
        metricsPort:
          $ref: '#/components/schemas/ServingPortMetricsPort'
        exposeExternally:
          $ref: '#/components/schemas/ServingPortExposeExternally'
        exposedUrl:
          $ref: '#/components/schemas/ServingPortExposedUrl'
        exposedProtocol:
          $ref: '#/components/schemas/ServingPortExposedProtocol'
      type:
      - object
      - 'null'
    CpuMemoryRequest:
      description: The amount of CPU memory to allocate for this workload (1G, 20M, .etc). The workload will receive at least this amount of memory. Note that the workload will not be scheduled unless the system can guarantee this amount of memory to the workload
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 20M
    GpuDevicesRequest:
      description: Requested number of GPU devices. Currently if more than one device is requested, it is not possible to provide values for gpuMemory or gpuPortion.
      type:
      - integer
      - 'null'
      format: int32
      example: 1
      minimum: 0
    TenantId:
      description: The id of the tenant.
      type: integer
      format: int32
      example: 1001
    PriorityClass:
      description: 'Specifies the priority class for the workload, which determines its scheduling behavior. Valid values are: very-low, low, medium-low, medium, medium-high, high, and very-high. Each workload type has a default priority. To view the default priority for each workload type, use the GET /workload-types endpoint. Once you change the priority from the default value defined for that workload type, the preemptibility field is not automatically updated. Make sure to set the desired preemptibility value.'
      type:
      - string
      - 'null'
      pattern: .*
    Label:
      description: Label details to be populated into the container.
      properties:
        name:
          description: The name of the label (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: stage
          pattern: .*
        value:
          description: The value of the label.
          type:
          - string
          - 'null'
          example: initial-research
          pattern: .*
        exclude:
          description: Use 'true' in case the label is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
      type:
      - object
      - 'null'
    NimServiceUpdateRequest:
      type: object
      required:
      - spec
      properties:
        spec:
          $ref: '#/components/schemas/NimServiceSpec'
    SubmissionErrorV2:
      allOf:
      - $ref: '#/components/schemas/Error'
      - $ref: '#/components/schemas/ComplianceIssuesV2'
    NimCache:
      description: The specification of a NIM cache volume.
      type:
      - object
      - 'null'
      properties:
        name:
          description: The NIMCache resource name (mandatory).
          type:
          - string
          - 'null'
          minLength: 1
          example: nim-cache-a
        profile:
          description: The NIM profile to use (optional).
          type:
          - string
          - 'null'
          minLength: 1
          example: tensorrt_llm-b200-fp8-tp2-pp1-latency-2901:10de-2
    ImagePullSecrets:
      description: A list of references to Kubernetes secrets in the same namespace used for pulling container images.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/ImagePullSecret'
    PvcVolumeMode:
      description: Default volume mode for the PVC. Choose between Filesystem (default) or Block.
      type:
      - string
      - 'null'
      enum:
      - Filesystem
      - Block
    ServingPortServiceType:
      description: The type of Kubernetes service to create for the inference deployment. Options include 'ClusterIP' (default), 'NodePort', 'LoadBalancer', and 'ExternalName'.
      type:
      - string
      - 'null'
      default: ClusterIP
      enum:
      - ClusterIP
      - NodePort
      - LoadBalancer
      - ExternalName
    ServingPortExposedProtocol:
      description: The protocol to use for the exposed URL. If grpcPort is set, this defaults to grpc. Otherwise, it defaults to http.
      type:
      - string
      - 'null'
      enum:
      - http
      - grpc
    TolerationEffect:
      description: The taint effect to match. (mandatory)
      type:
      - string
      - 'null'
      enum:
      - NoSchedule
      - NoExecute
      - PreferNoSchedule
      - Any
    Annotation:
      description: Annotation details to be populated into the container.
      properties:
        name:
          description: The name of the annotation (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          maxLength: 63
          example: billing
          pattern: .*
        value:
          description: The value of the annotation.
          type:
          - string
          - 'null'
          example: my-billing-unit
          pattern: .*
        exclude:
          description: Use 'true' in case the annotation is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          default: false
          example: false
      type:
      - object
      - 'null'
    PvcClaimSize:
      description: 'Requested size for the PVC. Mandatory when existingPvc is false. Recommended sizes: TB/GB/MB/TIB/GIB/MIB'
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 1G
    EnvironmentVariablePodFieldReference:
      description: Details of the field-reference and key use to populate the environment variable
      properties:
        path:
          description: The field path resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: metadata.name
          pattern: .*
      type:
      - object
      - 'null'
    Labels:
      description: Set of labels to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Label'
    DepartmentName1:
      type: string
      description: The name of the department
      example: default
      minLength: 1
      pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
    Preemptibility:
      description: Specifies whether the workload can be preempted by higher-priority workloads. Valid values are preemptible and non-preemptible. If explicitly set, this value takes precedence. If not set, the system derives the preemptibility from the priorityClassName field, ensuring backward compatibility. Each workload type has a default preemptibility. To view the default preemptibility for each workload type, use the GET /workload-types endpoint.
      type:
      - string
      - 'null'
      minLength: 1
      enum:
      - preemptible
      - non-preemptible
    PvcAddedAttrValues:
      description: an optional array of key-values pairs that are written as annotations on the created PVC. the allowed attributes are determined according to the storage class configuration (see k8s-objects-tracker for further info).
      type: array
      items:
        $ref: '#/components/schemas/PvcAddedAttrValue'
    RunAsUid:
      description: The user id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset `runAsUid` field (optional). Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled.
      type:
      - integer
      - 'null'
      format: int64
      example: 500
    NimServiceReplicas:
      default: 1
      description: The number of replicas to deploy.
      type:
      - integer
      - 'null'
      format: int32
      minimum: 0
      maximum: 1000
      example: 2
    AutoScalingMaxReplicas:
      description: The maximum number of replicas for autoscaling. Defaults to minReplicas. Must be no less than minReplicas.
      type:
      - integer
      - 'null'
      format: int32
      minimum: 1
    NimServiceSpec:
      allOf:
      - properties:
          annotations:
            $ref: '#/components/schemas/Annotations'
          autoscaling:
            properties:
              maxReplicas:
                $ref: '#/components/schemas/AutoScalingMaxReplicas'
              metric:
                $ref: '#/components/schemas/AutoScalingMetricNim'
              metricThreshold:
                $ref: '#/components/schemas/AutoScalingMetricThreshold'
              minReplicas:
                $ref: '#/components/schemas/AutoScalingMinReplicas'
              scaleWindowSeconds:
                $ref: '#/components/schemas/AutoScalingScaleWindowSeconds'
            type:
            - object
            - 'null'
          category:
            $ref: '#/components/schemas/Category'
          compute:
            properties:
              cpuCoreLimit:
                $ref: '#/components/schemas/CpuCoreLimit'
              cpuCoreRequest:
                $ref: '#/components/schemas/CpuCoreRequest'
              cpuMemoryLimit:
                $ref: '#/components/schemas/CpuMemoryLimit'
              cpuMemoryRequest:
                $ref: '#/components/schemas/CpuMemoryRequest'
              gpuDevicesRequest:
                $ref: '#/components/schemas/GpuDevicesRequest'
              gpuMemoryLimit:
                $ref: '#/components/schemas/GpuMemoryLimit'
              gpuMemoryRequest:
                $ref: '#/components/schemas/GpuMemoryRequest'
              gpuPortionLimit:
                $ref: '#/components/schemas/GpuPortionLimit'
              gpuPortionRequest:
                $ref: '#/components/schemas/GpuPortionRequest'
              gpuRequestType:
                $ref: '#/components/schemas/GpuRequestType'
            type:
            - object
            - 'null'
          environmentVariables:
            $ref: '#/components/schemas/EnvironmentVariables'
          image:
            $ref: '#/components/schemas/Image'
          imagePullPolicy:
            $ref: '#/components/schemas/ImagePullPolicy'
          imagePullSecrets:
            $ref: '#/components/schemas/ImagePullSecrets'
          labels:
            $ref: '#/components/schemas/Labels'
          modelStore:
            properties:
              nimCache:
                $ref: '#/components/schemas/NimCache'
              pvc:
                $ref: '#/components/schemas/NimServicePvcFields'
            type:
            - object
            - 'null'
          multiNode:
            $ref: '#/components/schemas/NimServiceMultiNode'
          ngcAuthSecret:
            $ref: '#/components/schemas/NimServiceNgcAuthSecret'
          nodePools:
            $ref: '#/components/schemas/NodePools'
          preemptibility:
            $ref: '#/components/schemas/Preemptibility'
          priorityClass:
            $ref: '#/components/schemas/PriorityClass'
          probes:
            $ref: '#/components/schemas/Probes'
          replicas:
            $ref: '#/components/schemas/NimServiceReplicas'
          security:
            properties:
              runAsGid:
                $ref: '#/components/schemas/RunAsGid'
              runAsUid:
                $ref: '#/components/schemas/RunAsUid'
            type:
            - object
            - 'null'
          servingPort:
            $ref: '#/components/schemas/NimServiceServingPort'
          tolerations:
            $ref: '#/components/schemas/Tolerations'
        type:
        - object
        - 'null'
    WorkloadName:
      description: The name of the workload.
      type: string
      minLength: 1
      example: my-workload-name
      pattern: .*
    Image:
      description: Docker image name. For more information, see [Images](https://kubernetes.io/docs/concepts/containers/images). The image name is mandatory for creating a workload.
      type:
      - string
      - 'null'
      minLength: 1
      example: python:3.8
      pattern: .*
    EnvironmentVariable:
      description: Details of an environment variable which is populated into the container.
      properties:
        name:
          description: The name of the environment variable. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: HOME
          pattern: .*
        value:
          description: The value of the environment variable. (mutually exclusive with secret, userCredential, configMap and podFieldRef)
          type:
          - string
          - 'null'
          example: /home/my-folder
          pattern: .*
        secret:
          $ref: '#/components/schemas/EnvironmentVariableSecret'
        configMap:
          $ref: '#/components/schemas/EnvironmentVariableConfigMap'
        podFieldRef:
          $ref: '#/components/schemas/EnvironmentVariablePodFieldReference'
        userCredential:
          $ref: '#/components/schemas/EnvironmentVariableUserCredential'
        exclude:
          description: Use 'true' in case the environment variable is defined in defaults of the policy, and you wish to exclude it from the workload.
          type:
          - boolean
          - 'null'
          example: false
        description:
          description: Description of the environment variable.
          type:
          - string
          - 'null'
          example: Home directory of the user.
          pattern: .*
      type:
      - object
      - 'null'
    PvcAccessModes:
      description: Default access mode(s) applied to newly created PVCs unless explicitly overridden.
      properties:
        readWriteOnce:
          description: Mount the volume as read/write by a single node.
          type:
          - boolean
          - 'null'
          default: true
        readOnlyMany:
          description: Mount the volume as read-only by many nodes.
          type:
          - boolean
          - 'null'
          default: false
        readWriteMany:
          description: Mount the volume as read/write by many nodes.
          type:
          - boolean
          - 'null'
          default: false
      type:
      - object
      - 'null'
    EnvironmentVariables:
      description: Set of environment variables to populate into the container running the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/EnvironmentVariable'
    Probe:
      type:
      - object
      - 'null'
      properties:
        initialDelaySeconds:
          description: Number of seconds after the container has started before liveness or readiness probes are initiated.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 0
        periodSeconds:
          description: How often (in seconds) to perform the probe.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        timeoutSeconds:
          description: Number of seconds after which the probe times out.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        successThreshold:
          description: Minimum consecutive successes for the probe to be considered successful after having failed.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        failureThreshold:
          description: When a probe fails, the number of times to try before giving up.
          type:
          - integer
          - 'null'
          format: int32
          minimum: 1
        handler:
          $ref: '#/components/schemas/ProbeHandler'
    ServingPortExposeExternally:
      description: Indicates whether the inference serving endpoint should be accessible outside the cluster. If set to true, the endpoint will be exposed externally. To enable external access, your administrator must configure the cluster as described in the [inference requirements](https://run-ai-docs.nvidia.com/saas/getting-started/installation/system-requirements#inference). section.
      type:
      - boolean
      - 'null'
      default: true
    ProjectId:
      description: The id of the project.
      type: string
      example: 1
      pattern: .*
    Tolerations:
      description: Set of tolerations to apply to the workload.
      type:
      - array
      - 'null'
      items:
        $ref: '#/components/schemas/Toleration'
    NIMServiceMetadataCreateParams:
      type: object
      required:
      - name
      - projectId
      properties:
        name:
          $ref: '#/components/schemas/WorkloadName'
        useGivenNameAsPrefix:
          description: When true, the requested name will be treated as a prefix. The final name of the workload will be composed of the name followed by a random set of characters.
          type: boolean
          example: true
          default: false
        projectId:
          $ref: '#/components/schemas/ProjectId'
    ComplianceIssuesV2:
      properties:
        complianceIssues:
          type: array
          items:
            type: object
            required:
            - details
            - field
            properties:
              field:
                type: string
                example: compute.gpuDevicesRequest
              details:
                type: string
                example: value must be no less than 3
              rule:
                $ref: '#/components/schemas/PolicyRuleEnum'
      type:
      - object
      - 'null'
    EnvironmentVariableSecret:
      description: Details of the secret and key use to populate the environment variable
      properties:
        name:
          description: The name of the secret resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: postgress_secret
          pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
        key:
          description: The key in the secret resource. (mandatory)
          type:
          - string
          - 'null'
          minLength: 1
          example: POSTGRES_PASSWORD
          pattern: .*
      type:
      - object
      - 'null'
    GVK:
      type: object
      description: Specifies the Group, Version, and Kind (GVK) of the Kubernetes resource that defines the workload.
      required:
      - group
      - version
      - kind
      properties:
        group:
          description: The API group of the Kubernetes resource.
          type: string
          example: apps
        version:
          description: The API version of the resource within the specified group.
          type: string
          example: v1
        kind:
          description: The type of Kubernetes resource being referenced.
          type: string
          example: Deployment
    Category:
      description: Specify the workload category assigned to the workload. Categories are used to classify and monitor different types of workloads within the NVIDIA Run:ai platform.
      type:
      - string
      - 'null'
      pattern: .*
    PolicyRuleEnum:
      description: Indicates which validation rule (e.g., min, max, step, options, required, canEdit, canAdd) conflicted with policy restrictions, causing the asset or template to be rejected.
      type:
      - string
      - 'null'
      enum:
      - min
      - max
      - step
      - options
      - required
      - canEdit
      - canAdd
      - locked
    NimServiceResponse:
      type: object
      required:
      - spec
      - metadata
      - desiredPhase
      properties:
        metadata:
          $ref: '#/components/schemas/WorkloadV2Metadata'
        desiredPhase:
          $ref: '#/components/schemas/DesiredPhase'
        spec:
          $ref: '#/components/schemas/NimServiceSpec'
    GpuMemoryRequest:
      description: Required if and only if gpuRequestType is memory. States the GPU memory to allocate for the created workload, per GPU device. Note that the workload will not be scheduled unless the system can guarantee this amount of GPU memory to the workload.
      type:
      - string
      - 'null'
      pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
      example: 10M
    ProbeHandlerScheme:
      description: Scheme to use for connecting to the host, defaults to HTTP.
      type:
      - string
      - 'null'
      enum:
      - HTTP
      - HTTPS
    NodePools:
      description: A prioritized list of node pools for the scheduler to run the workload on. The scheduler will always try to use the first node pool before moving to the next one if the first is not available.
      type:
      - array
      - 'null'
      items:
        type: string
        pattern: .*
      example:
      - my-node-pool-a
      - my-node-pool-b
    RunAsGid:
      description: The group id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset `runAsGid` field (optional). Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled.
      type:
      - integer
      - 'null'
      format: int64
      example: 30
    ProbeHandler:
      description: The action taken to determine the health of the container. (mandatory)
      type:
      - object
      - 'null'
      properties:
        httpGet:
          description: An action based on HTTP Get requests.
          type: object
          properties:
            path:
              description: Path to access on the HTTP server, defaults to /.
              type:
              - string
              - 'null'
              pattern: ^(\x2F[a-zA-Z0-9\-_.\x2F]*)?$
              example: /
            port:
              description: Number of the port to access on the container.
              type:
              - integer
              - 'null'
              format: int32
              minimum: 1
              maximum: 65535
            host:
              description: Host name to connect to, defaults to the pod IP.
              type:
              - string
              - 'null'
              format: hostname
              example: example.com
              pattern: .*
            scheme:
              $ref: '#/components/schemas/ProbeHandlerScheme'
    WorkloadV2MetadataAutoFill:
      type: object
      required:
      - id
      - gvk
      - projectName
      - clusterId
      - tenantId
      - departmentId
      - departmentName
      - createdAt
      - createdBy
      - updatedAt
      - updatedBy
      properties:
        id:
          $ref: '#/components/schemas/WorkloadId3'
        gvk:
          $ref: '#/components/schemas/GVK'
        projectName:
          $ref: '#/components/schemas/ProjectName2'
        clusterId:
          $ref: '#/components/schemas/ClusterId'
        tenantId:
          $ref: '#/components/schemas/TenantId'
        departmentId:
          $ref: '#/components/schemas/DepartmentId3'
        departmentName:
          $ref: '#/components/schemas/DepartmentName1'
        createdAt:
          type: string
          format: date-time
          description: The timestamp for when the workload was created.
          example: '2024-01-15T10:30:00Z'
        createdBy:
          type: string
          description: Identifier of the user who created the workload.
          format: .*
          example: user@run.ai
        updatedAt:
          type: string
          format: date-time
          description: The timestamp for the last time the workload was updated.
          example: '2024-01-15T10:35:00Z'
        updatedBy:
          type: string
          description: Identifier of the user who last updated the workload.
          format: .*
          example: user@run.ai
        deletedAt:
          type:
          - string
          - 'null'
          format: date-time
          description: The timestamp indicating when the workload was deleted.
          example: '2024-01-15T10:35:00Z'
        deletedBy:
          type:
          - string
          - 'null'
          format: .*
          description: Identifier of the user who deleted the workload.
          example: user@run.ai
    ClaimInfo:
      description: Claim information for the newly created PVC. The information should not be provided when attempting to use existing PVC.
      properties:
        size:
          $ref: '#/components/schemas/PvcClaimSize'
        storageClass:
          description: Storage class name to associate with the PVC. This parameter may be omitted if there is a single storage class in the system, or you are using the default storage class. For more information, see [Storage class](https://kubernetes.io/docs/concepts/storage/storage-classes).
          type:
          - string
          - 'null'
          minLength: 1
          example: my-storage-class
          pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
        accessModes:
          $ref: '#/components/schemas/PvcAccessModes'
        volumeMode:
          $ref: '#/components/schemas/PvcVolumeMode'
        addedAttrValues:
          $ref: '#/components/schemas/PvcAdded

# --- truncated at 32 KB (46 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/runai/refs/heads/main/openapi/runai-nvidia-nim-api-openapi.yml