Every API here is available over the APIs.io API and to AI agents over MCP.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for apis
7 MCP tools reach this
find_apisBrowse and filter every API in the catalog.
get_api_artifactsOne API's artifacts, grouped by type.
get_openapiThe primary OpenAPI for this API.
find_similar_apisAPIs that look like this one.
apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
resolveTurn a domain, URL or GitHub org into the provider it belongs to.
find_cohortsEvery scored population of providers in the catalog.
All 92 tools →
Call it yourself
curl for this page
This API
curl "https://apis.io/api/v1/apis/runai-distributed-inferences-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.
Get an API key
Free tier, no form to fill in. Signing in shares your email address with us — we
store it to create your key and to recognise you if you sign in with another
provider. See our Privacy Policy and
Terms.
A second provider on the same verified email joins the account you already have.
openapi: 3.2.0
info:
version: latest
description: '# Introduction
The NVIDIA Run:ai Control-Plane API reference is a guide that provides an easy-to-use programming interface for adding various tasks to your application, including workload submission, resource management, and administrative operations.
NVIDIA Run:ai APIs are accessed using *bearer tokens*. To obtain a token, you need to create a **Service account** through the NVIDIA Run:ai user interface.
To create a service account, in your UI, go to Access → Service Accounts (for organization-level service accounts) or User settings → Access Keys (for user access keys), and create a new one.
After you have created a new service account, you will need to assign it access rules.
To assign access rules to the service account, see [Create access rules](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/accessrules#create-or-delete-rules).
Make sure you assign the correct rules to your service account. Use the [Roles](https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/roles) to assign the correct access rules.
To get your access token, follow the instructions in [Request a token](https://run-ai-docs.nvidia.com/saas/reference/api/rest-auth/#request-an-api-token).
'
title: NVIDIA Run:ai Access Keys Distributed Inferences API
x-logo:
url: https://api.redocly.com/registry/raw/runai-xq8/saas/latest/public/runai-logo-api.png
altText: NVIDIA Run:ai
href: https://run.ai
license:
name: NVIDIA Run:ai
url: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-software-license-agreement/
servers:
- url: https://app.run.ai
security:
- bearerAuth: []
tags:
- name: Distributed Inferences
description: "Distributed inference enables running inference workloads across multiple pods, typically to scale model serving beyond a single container or node. This approach is useful when a single instance cannot meet resource requirements.NVIDIA Run:ai supports this model using Leader Worker Set (LWS). \nEach pod plays a specific role, either as a leader or worker, and together they form a coordinated service. NVIDIA Run:ai manages the orchestration and configuration of these pods to ensure efficient and scalable inference execution\n"
paths:
/api/v1/workloads/distributed-inferences:
post:
summary: Create a distributed inference. [Experimental]
operationId: create_distributed_inference
description: Create a distributed inference using container related fields.
tags:
- Distributed Inferences
requestBody:
content:
application/json:
schema:
$ref: '#/components/schemas/DistributedInferenceCreationRequest'
responses:
'202':
description: Request completed successfully.
content:
application/json:
schema:
$ref: '#/components/schemas/DistributedInference'
'400':
$ref: '#/components/responses/400BadRequest'
'401':
$ref: '#/components/responses/401Unauthorized'
'403':
$ref: '#/components/responses/403Forbidden'
'503':
$ref: '#/components/responses/503ServiceUnavailable'
/api/v1/workloads/distributed-inferences/{workloadId}:
delete:
summary: Delete a distributed inference.
operationId: delete_distributed_inference
description: Delete a distributed inference using a workload id.
tags:
- Distributed Inferences
parameters:
- $ref: '#/components/parameters/WorkloadId'
responses:
'202':
$ref: '#/components/responses/202Accepted'
'401':
$ref: '#/components/responses/401Unauthorized'
'403':
$ref: '#/components/responses/403Forbidden'
'404':
$ref: '#/components/responses/404NotFound'
'500':
$ref: '#/components/responses/500InternalServerError'
'503':
$ref: '#/components/responses/503ServiceUnavailable'
get:
summary: Get a distributed inference data.
operationId: get_distributed_inference
description: Retrieve a distributed inference details using a workload id.
tags:
- Distributed Inferences
parameters:
- $ref: '#/components/parameters/WorkloadId'
responses:
'200':
description: Executed successfully.
content:
application/json:
schema:
$ref: '#/components/schemas/DistributedInference'
'401':
$ref: '#/components/responses/401Unauthorized'
'403':
$ref: '#/components/responses/403Forbidden'
'404':
$ref: '#/components/responses/404NotFound'
'500':
$ref: '#/components/responses/500InternalServerError'
'503':
$ref: '#/components/responses/503ServiceUnavailable'
patch:
summary: Update distributed inference spec.
operationId: update_distributed_inference_spec
description: Update the specification of an existing distributed inference workload.
tags:
- Distributed Inferences
parameters:
- $ref: '#/components/parameters/WorkloadId'
requestBody:
content:
application/json:
schema:
$ref: '#/components/schemas/UpdateRequest'
responses:
'202':
description: Executed successfully.
content:
application/json:
schema:
$ref: '#/components/schemas/DistributedInference'
'401':
$ref: '#/components/responses/401Unauthorized'
'403':
$ref: '#/components/responses/403Forbidden'
'404':
$ref: '#/components/responses/404NotFound'
'500':
$ref: '#/components/responses/500InternalServerError'
'503':
$ref: '#/components/responses/503ServiceUnavailable'
components:
schemas:
LargeShmRequest:
description: A large /dev/shm device to mount into a container running the created workload. An shm is a shared file system mounted on RAM.
type:
- boolean
- 'null'
example: false
SecretFieldsUpdatable:
properties:
mountPath:
description: Local path within the workload to which the Secret will be mapped to. (mandatory)
type:
- string
- 'null'
minLength: 1
defaultMode:
$ref: '#/components/schemas/DefaultMode'
type:
- object
- 'null'
EnvironmentVariableConfigMap:
description: Details of the configMap and key use to populate the environment variable
properties:
name:
description: The name of the config-map resource. (mandatory)
type:
- string
- 'null'
minLength: 1
example: my-config-map
pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
key:
description: The key in the config-map resource. (mandatory)
type:
- string
- 'null'
minLength: 1
example: MY_POSTGRES_SCHEMA
pattern: .*
type:
- object
- 'null'
ConfigMapInstance:
allOf:
- $ref: '#/components/schemas/StorageInstanceName'
- $ref: '#/components/schemas/ConfigMap'
- $ref: '#/components/schemas/ExcludeField'
type:
- object
- 'null'
WorkloadId2:
description: A unique ID of the workload.
type: string
format: uuid
DefaultMode:
type:
- string
- 'null'
description: 'File permission mode in octal string format. This value must be a 4-digit octal number, representing the default file mode when mounting a Secret or ConfigMap as a volume.
'
minLength: 4
maxLength: 4
example: '0644'
pattern: 0[0-7]{3}
CpuMemoryRequest:
description: The amount of CPU memory to allocate for this workload (1G, 20M, .etc). The workload will receive at least this amount of memory. Note that the workload will not be scheduled unless the system can guarantee this amount of memory to the workload
type:
- string
- 'null'
pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
example: 20M
GpuDevicesRequest:
description: Requested number of GPU devices. Currently if more than one device is requested, it is not possible to provide values for gpuMemory or gpuPortion.
type:
- integer
- 'null'
format: int32
example: 1
minimum: 0
PriorityClass:
description: 'Specifies the priority class for the workload, which determines its scheduling behavior. Valid values are: very-low, low, medium-low, medium, medium-high, high, and very-high. Each workload type has a default priority. To view the default priority for each workload type, use the GET /workload-types endpoint. Once you change the priority from the default value defined for that workload type, the preemptibility field is not automatically updated. Make sure to set the desired preemptibility value.'
type:
- string
- 'null'
pattern: .*
Label:
description: Label details to be populated into the container.
properties:
name:
description: The name of the label (mandatory)
type:
- string
- 'null'
minLength: 1
maxLength: 63
example: stage
pattern: .*
value:
description: The value of the label.
type:
- string
- 'null'
example: initial-research
pattern: .*
exclude:
description: Use 'true' in case the label is defined in defaults of the policy, and you wish to exclude it from the workload.
type:
- boolean
- 'null'
example: false
type:
- object
- 'null'
DepartmentId2:
description: The id of the department.
type: string
minLength: 1
example: 2
pattern: .*
DistributedInferenceStartupPolicyField:
properties:
startupPolicy:
$ref: '#/components/schemas/DistributedInferenceStartupPolicy'
type:
- object
- 'null'
NodeSelectorTerm:
type:
- object
- 'null'
description: A null or empty node selector term matches no objects. The requirements of them are ANDed.
properties:
matchExpressions:
description: A list of node selector requirements by node's labels.
type: array
items:
$ref: '#/components/schemas/MatchExpression'
ImagePullSecrets:
description: A list of references to Kubernetes secrets in the same namespace used for pulling container images.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/ImagePullSecret'
PvcVolumeMode:
description: Default volume mode for the PVC. Choose between Filesystem (default) or Block.
type:
- string
- 'null'
enum:
- Filesystem
- Block
TolerationEffect:
description: The taint effect to match. (mandatory)
type:
- string
- 'null'
enum:
- NoSchedule
- NoExecute
- PreferNoSchedule
- Any
Annotation:
description: Annotation details to be populated into the container.
properties:
name:
description: The name of the annotation (mandatory)
type:
- string
- 'null'
minLength: 1
maxLength: 63
example: billing
pattern: .*
value:
description: The value of the annotation.
type:
- string
- 'null'
example: my-billing-unit
pattern: .*
exclude:
description: Use 'true' in case the annotation is defined in defaults of the policy, and you wish to exclude it from the workload.
type:
- boolean
- 'null'
default: false
example: false
type:
- object
- 'null'
NodeType3:
description: Nodes (machines), or a group of nodes on which the workload will run. To use this feature, your Administrator will need to label nodes. For more information, see [Group Nodes](https://docs.run.ai/latest/admin/researcher-setup/limit-to-node-group). When using this flag with with Project-based affinity, it refines the list of allowable node groups set in the Project. For more information, see [Projects](https://docshub.run.ai/guides/platform-management/aiinitiatives/organization/projects).
type:
- string
- 'null'
minLength: 1
example: my-node-type
pattern: .*
PvcClaimSize:
description: 'Requested size for the PVC. Mandatory when existingPvc is false. Recommended sizes: TB/GB/MB/TIB/GIB/MIB'
type:
- string
- 'null'
pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
example: 1G
EnvironmentVariablePodFieldReference:
description: Details of the field-reference and key use to populate the environment variable
properties:
path:
description: The field path resource. (mandatory)
type:
- string
- 'null'
minLength: 1
example: metadata.name
pattern: .*
type:
- object
- 'null'
Labels:
description: Set of labels to populate into the container running the workload.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/Label'
PvcFieldsNonUpdatable:
properties:
existingPvc:
description: Verify existing PVC. PVC is assumed to exist when set to `true`. If set to `false`, the PVC will be created, if it does not exist.
type:
- boolean
- 'null'
default: false
claimName:
description: Name for the PVC. Allow referencing it across workloads. If not provided, a name based on the workload name and scope will be auto-generated.
type:
- string
- 'null'
minLength: 1
maxLength: 63
example: my-claim
pattern: .*
readOnly:
description: Permit only read access to PVC.
type:
- boolean
- 'null'
default: false
ephemeral:
description: Use `true` to set PVC to ephemeral. If set to `true`, the PVC will be deleted when the workload is stopped. Not supported for inference workloads.
type:
- boolean
- 'null'
default: false
example: false
claimInfo:
$ref: '#/components/schemas/ClaimInfo'
dataSharing:
description: use `true` to share the PVC data to all projects under the selected scope.
type:
- boolean
- 'null'
default: false
example: false
type:
- object
- 'null'
DistributedInferenceSpecSpec:
allOf:
- $ref: '#/components/schemas/DistributedInferenceCommonSpec'
- $ref: '#/components/schemas/DistributedInferenceLeaderSpecFields'
- $ref: '#/components/schemas/DistributedInferenceWorkerSpecFields'
Preemptibility:
description: Specifies whether the workload can be preempted by higher-priority workloads. Valid values are preemptible and non-preemptible. If explicitly set, this value takes precedence. If not set, the system derives the preemptibility from the priorityClassName field, ensuring backward compatibility. Each workload type has a default preemptibility. To view the default preemptibility for each workload type, use the GET /workload-types endpoint.
type:
- string
- 'null'
minLength: 1
enum:
- preemptible
- non-preemptible
PvcAddedAttrValues:
description: an optional array of key-values pairs that are written as annotations on the created PVC. the allowed attributes are determined according to the storage class configuration (see k8s-objects-tracker for further info).
type: array
items:
$ref: '#/components/schemas/PvcAddedAttrValue'
RunAsUid:
description: The user id to run the entrypoint of the container which executes the workspace. Default to the value specified in the environment asset `runAsUid` field (optional). Use only when the source uid/gid of the environment asset is not `fromTheImage`, and `overrideUidGidInWorkspace` is enabled.
type:
- integer
- 'null'
format: int64
example: 500
StorageInstanceName:
properties:
name:
description: unique name to identify the instance. primarily used for policy locked rules.
type:
- string
- 'null'
minLength: 1
example: storage-instance-a
type:
- object
- 'null'
WorkloadCreationMeta:
required:
- name
- projectId
- clusterId
properties:
name:
$ref: '#/components/schemas/WorkloadName'
useGivenNameAsPrefix:
description: When true, the requested name will be treated as a prefix. The final name of the workload will be composed of the name followed by a random set of characters.
type: boolean
example: true
default: false
projectId:
$ref: '#/components/schemas/ProjectId'
clusterId:
$ref: '#/components/schemas/ClusterId'
WorkloadName:
description: The name of the workload.
type: string
minLength: 1
example: my-workload-name
pattern: .*
Image:
description: Docker image name. For more information, see [Images](https://kubernetes.io/docs/concepts/containers/images). The image name is mandatory for creating a workload.
type:
- string
- 'null'
minLength: 1
example: python:3.8
pattern: .*
SeccompProfileType:
description: Indicates which kind of seccomp profile will be applied to the container. The options are a. `RuntimeDefault` - the container runtime default profile should be used. b. `Unconfined` - no profile should be applied. c. `Localhost` is not yet supported by Run:ai.
type:
- string
- 'null'
enum:
- RuntimeDefault
- Unconfined
- Localhost
EnvironmentVariable:
description: Details of an environment variable which is populated into the container.
properties:
name:
description: The name of the environment variable. (mandatory)
type:
- string
- 'null'
minLength: 1
example: HOME
pattern: .*
value:
description: The value of the environment variable. (mutually exclusive with secret, userCredential, configMap and podFieldRef)
type:
- string
- 'null'
example: /home/my-folder
pattern: .*
secret:
$ref: '#/components/schemas/EnvironmentVariableSecret'
configMap:
$ref: '#/components/schemas/EnvironmentVariableConfigMap'
podFieldRef:
$ref: '#/components/schemas/EnvironmentVariablePodFieldReference'
userCredential:
$ref: '#/components/schemas/EnvironmentVariableUserCredential'
exclude:
description: Use 'true' in case the environment variable is defined in defaults of the policy, and you wish to exclude it from the workload.
type:
- boolean
- 'null'
example: false
description:
description: Description of the environment variable.
type:
- string
- 'null'
example: Home directory of the user.
pattern: .*
type:
- object
- 'null'
DistributedInferenceStartupPolicy:
description: "Determines when the worker pods should start during workload initialization. \n - `LeaderCreated`: Workers start after the leader pod is created.\n - `LeaderReady`: Workers start only after the leader pod is ready.\n"
type:
- string
- 'null'
enum:
- LeaderCreated
- LeaderReady
default: LeaderCreated
PvcAccessModes:
description: Default access mode(s) applied to newly created PVCs unless explicitly overridden.
properties:
readWriteOnce:
description: Mount the volume as read/write by a single node.
type:
- boolean
- 'null'
default: true
readOnlyMany:
description: Mount the volume as read-only by many nodes.
type:
- boolean
- 'null'
default: false
readWriteMany:
description: Mount the volume as read/write by many nodes.
type:
- boolean
- 'null'
default: false
type:
- object
- 'null'
EnvironmentVariables:
description: Set of environment variables to populate into the container running the workload.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/EnvironmentVariable'
Probe:
type:
- object
- 'null'
properties:
initialDelaySeconds:
description: Number of seconds after the container has started before liveness or readiness probes are initiated.
type:
- integer
- 'null'
format: int32
minimum: 0
periodSeconds:
description: How often (in seconds) to perform the probe.
type:
- integer
- 'null'
format: int32
minimum: 1
timeoutSeconds:
description: Number of seconds after which the probe times out.
type:
- integer
- 'null'
format: int32
minimum: 1
successThreshold:
description: Minimum consecutive successes for the probe to be considered successful after having failed.
type:
- integer
- 'null'
format: int32
minimum: 1
failureThreshold:
description: When a probe fails, the number of times to try before giving up.
type:
- integer
- 'null'
format: int32
minimum: 1
handler:
$ref: '#/components/schemas/ProbeHandler'
PodAffinity:
description: Pod affinity scheduling rules (e.g. co-locate this workload in the same node, zone, etc. as some other workloads).
type:
- object
- 'null'
properties:
type:
$ref: '#/components/schemas/PodAffinityType'
key:
description: The label key to use. (mandatory)
type:
- string
- 'null'
pattern: .*
CreateHomeDir:
description: When set to `true`, creates a home directory for the container.
type:
- boolean
- 'null'
UpdateSpec:
description: The specifications of the inference to be updated.
properties:
spec:
allOf:
- type:
- object
- 'null'
properties:
replicas:
description: Specifies the number of leader-worker sets to deploy. Each replica represents a group consisting of one leader pod and multiple worker pods. For example, setting replicas to 3 will create 3 independent groups, each with its own leader and corresponding set of workers.
type:
- integer
- 'null'
format: int32
minimum: 0
maximum: 1000
example: 2
ProjectId:
description: The id of the project.
type: string
example: 1
pattern: .*
DistributedInferenceReplicasField:
properties:
replicas:
default: 1
description: "Specifies the number of leader-worker sets to deploy. Each replica represents a group consisting of one leader pod and multiple worker pods. \nFor example, setting replicas: 3 will create 3 independent groups, each with its own leader and corresponding set of workers.\n"
type:
- integer
- 'null'
format: int32
minimum: 0
maximum: 1000
example: 2
Tolerations:
description: Set of tolerations to apply to the workload.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/Toleration'
DistributedInference:
allOf:
- $ref: '#/components/schemas/WorkloadMeta1'
- $ref: '#/components/schemas/DistributedInferenceSpec'
ExtendedResource:
description: Quantity of an extended resource.
properties:
resource:
description: The name of the extended resource (mandatory)
type:
- string
- 'null'
example: hardware-vendor.example/foo
minLength: 1
pattern: .*
quantity:
description: The requested quantity for the resource.
type:
- string
- 'null'
example: 2
minLength: 1
pattern: ^([+]?[0-9.]+)([eEinumkKMGTP]*[-+]?[0-9]*)$
exclude:
description: Use 'true' in case the extended resource is defined in defaults of the policy, and you wish to exclude it from the workload.
type:
- boolean
- 'null'
example: false
type:
- object
- 'null'
DistributedInferenceServingPort:
description: Defines the configuration for the inference serving endpoint. This determines how applications or services can send inference requests to the workload.
allOf:
- $ref: '#/components/schemas/DistributedInferenceServingPortContainerAndProtocol'
- $ref: '#/components/schemas/DistributedInferenceServingPortAccess'
type:
- object
- 'null'
DistributedInferenceServingPortProtocol:
description: The protocol used to access the port.
type:
- string
- 'null'
enum:
- http
default: http
PvcItems:
description: Set of pvc persistent volume claims to use in the workload.
type:
- array
- 'null'
items:
$ref: '#/components/schemas/PvcInstance'
DistributedInferenceLeaderWorkerSpec1:
properties:
annotations:
$ref: '#/components/schemas/Annotations'
args:
$ref: '#/components/schemas/Args'
command:
$ref: '#/components/schemas/Command'
compute:
properties:
cpuCoreLimit:
$ref: '#/components/schemas/CpuCoreLimit'
cpuCoreRequest:
$ref: '#/components/schemas/CpuCoreRequest'
cpuMemoryLimit:
$ref: '#/components/schemas/CpuMemoryLimit'
cpuMemoryRequest:
$ref: '#/components/schemas/CpuMemoryRequest'
extendedResources:
$ref: '#/components/schemas/ExtendedResources'
gpuDevicesRequest:
$ref: '#/components/schemas/GpuDevicesRequest'
gpuMemoryLimit:
$ref: '#/components/schemas/GpuMemoryLimit'
gpuMemoryRequest:
$ref: '#/components/schemas/GpuMemoryRequest'
gpuPortionLimit:
$ref: '#/components/schemas/GpuPortionLimit'
gpuPortionRequest:
$ref: '#/components/schemas/GpuPortionRequest'
gpuRequestType:
$ref: '#/components/schemas/GpuRequestType'
largeShmRequest:
$ref: '#/components/schemas/LargeShmRequest'
type:
- object
- 'null'
createHomeDir:
$ref: '#/components/schemas/CreateHomeDir'
environmentVariables:
$ref: '#/components/schemas/EnvironmentVariables'
image:
$ref: '#/components/schemas/Image'
imagePullPolicy:
$ref: '#/components/schemas/ImagePullPolicy'
imagePullSecrets:
$ref: '#/components/schemas/ImagePullSecrets'
labels:
$ref: '#/components/schemas/Labels'
nodeAffinityRequired:
$ref: '#/components/schemas/NodeAffinityRequired'
nodeType:
$ref: '#/components/schemas/NodeType3'
podAffinity:
$ref: '#/components/schemas/PodAffinity'
probes:
$ref: '#/components/schemas/Probes'
security:
properties:
capabilities:
$ref: '#/components/schemas/Capabilities'
readOnlyRootFilesystem:
$ref: '#/components/schemas/ReadOnlyRootFileSystem'
runAsGid:
$ref: '#/components/schemas/RunAsGid'
runAsNonRoot:
$ref: '#/components/schemas/RunAsNonRoot'
runAsUid:
$ref: '#/components/schemas/RunAsUid'
seccompProfileType:
$ref: '#/components/schemas/SeccompProfileType'
supplementalGroups:
$ref: '#/components/schemas/SupplementalGroups'
uidGidSource:
$ref: '#/components/schemas/UidGidSource'
type:
- object
- 'null'
storage:
properties:
configMapVolume:
$ref: '#/components/schemas/ConfigMapItems'
emptyDirVolume:
$ref: '#/components/schemas/EmptyDirItems'
pvc:
$ref: '#/components/schemas/PvcItems'
secretVolume:
$ref: '#/components/schemas/SecretItems1'
type:
- object
- 'null'
tolerations:
$ref: '#/components/schemas/Tolerations'
workingDir:
$ref: '#/components/schemas/WorkingDir'
type: object
Args:
description: Arguments to the command that the container running the workload executes.
type:
- string
- 'null'
minLength: 1
example: -x my-script.py
pattern: .*
DistributedInferenceServingPortAccess:
properties:
authorizationType:
$ref: '#/components/schemas/DistributedInferenceServingPortAccessAuthorizationTypeEnum'
authorizedUsers:
description: A list of users and service accounts allowed to send inference requests to the serving endpoint. `Note:` Cannot be used together with authorizedGroups.
type:
- array
- 'null'
items:
type: string
pattern: .*
example:
- user.a@example.com
- user.b@example.com
authorizedGroups:
description: A list of user groups allowed to send inference requests to the serving endpoint. `Note:` Cannot be used together with authorizedUsers.
type:
- array
- 'null'
items:
type: string
pattern: .*
example:
- group-a
- group-b
exposeExternally:
description: Indicates whether the inference serving endpoint should be accessible outside the cluster. If set to true, the endpoint will be exposed externally. To enable external access, your administrator must configure the cluster as described in the [inference requirements](https://run-ai-docs.nvidia.com/saas/getting-started/installation/system-requirements#inference). section.
type:
- boolean
- 'null'
default: true
exposedUrl:
description: The custom URL to use for the serving port. If empty (default), an autogenerated URL will be used.
type:
- string
- 'null'
pattern: .*
EnvironmentVariableSecret:
description: Details of the secret and key use to populate the environment variable
properties:
name:
description: The name of the secret resource. (mandatory)
type:
- string
- 'null'
minLength: 1
example: postgress_secret
pattern: ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$
key:
description: The key in the secret resource. (mandatory)
type:
- string
- 'null'
minLength: 1
example: POSTGRES_PASSWORD
pattern: .*
type:
- object
- 'null'
Command:
description: A command to the server as the entry point of the container running the workload.
type:
- string
- 'null'
minLength: 1
example: python
pattern: .
# --- truncated at 32 KB (61 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/runai/refs/heads/main/openapi/runai-distributed-inferences-api-openapi.yml