Helm Chart Changes in Version 8.0.0
Version 8.0.0 of Capacity Private Cloud introduces new services, GPU acceleration, a new way of selecting ASR models, and a new optional ingress mode. Most of these are controlled through your Helm values file. This article explains each new setting, shows example values, and lists the changes you need to make when upgrading from 7.x.
Applies to: Capacity Private Cloud 8.0.0 Speech products (charts lumenvox, lumenvox-common, lumenvox-speech and lumenvox-external-services 8.0.0).
For the full list of changes in this release, see the Release Notes 8.0.0.
Before you begin
Kubernetes 1.35 or later is required. Version 7.x required 1.33.
Update your local Helm chart repository before installing or upgrading:
helm repo updateThe default image tag is now
:8.0. If your values file does not setglobal.image.tag, your services will move to 8.0 images, and:8.0always pulls the latest 8.0.x patch.Review your values file before upgrading. Some 7.x settings have been renamed or moved and will stop the install with a message explaining what to change. See Upgrading from 7.x below.
Summary of new settings
| Feature | Helm key | Default |
|---|---|---|
| GPU inference (per service) | global.gpu.<service>.enabled | Off |
| GPU per language | global.asrLanguages[].services.<service>.gpu | Uses the global default |
| ASR service | global.asr.enabled | On |
| Real-time transcription (new) | global.transcribeRealtime.enabled | On |
| Batch transcription (new) | global.transcribeBatch.enabled | On |
| Neural TTS | global.neuralTts.enabled | On |
| Legacy TTS for a language | global.ttsLanguages[].legacyEnabled | Off (Neural TTS is used) |
| Additional ASR models for a language | global.asrLanguages[].extraModels | None |
| Default ASR model version | global.asrDefaultVersion | "8.0.0" |
| Istio Gateway API ingress (new) | global.lumenvox.ingress.className: "istio" | "nginx" |
| In-cluster databases for test/dev (new) | global.enabled.externalServices | Off |
Throughout this article, <service> means asr, transcribeRealtime or transcribeBatch.
1. GPU acceleration for ASR and transcription
ASR, real-time transcription and batch transcription can each run with GPU inference. GPU processing reduces latency and allows additional throughput. It is off by default.
You set a chart-wide default for each service, and can then override it for individual languages.
Turn on GPU for all three services, for every language:
global:
gpu:
asr:
enabled: true
count: 1 # nvidia.com/gpu devices per pod
transcribeRealtime:
enabled: true
transcribeBatch:
enabled: true
Or turn on GPU per language (in this example, GPU for English only, with batch transcription using two GPUs):
global:
asrLanguages:
- name: "en"
services:
asr:
gpu: true
transcribeRealtime:
gpu: true
transcribeBatch:
gpu: { enabled: true, count: 2, visibleDevices: "all" }
- name: "es" # no gpu block: uses the global.gpu defaults
Available GPU options
| Option | Description |
|---|---|
enabled | Turns GPU inference on or off. |
count | Number of nvidia.com/gpu devices per pod. Default 1. |
visibleDevices | Sets NVIDIA_VISIBLE_DEVICES (for example "all"). |
runtimeClassName | Container runtime class to use (for example "nvidia"). |
At the language level, a plain true or false overrides only enabled. An object (as in the transcribeBatch example above) can override any of the four options.
Requirements and notes
- GPU nodes with the NVIDIA device plugin installed.
- GPU mode is currently available only with high-definition acoustic models. See section 5 for how to add high-definition models to a language.
- GPU pods automatically tolerate the
nvidia.com/gputaint and the AKS spot-node taint, so they schedule onto tainted GPU node pools without extra settings. - Set
visibleDevices: "all"on clusters where the device plugin setsNVIDIA_VISIBLE_DEVICES=void. - Set
runtimeClassNamewhere the node's default container runtime is not the NVIDIA runtime.
2. Turning speech services on and off
Each speech service can now be switched on or off independently. All four are on by default.
- ASR, real-time transcription and batch transcription are deployed once for each entry in
asrLanguages. - Neural TTS is deployed once for each entry in
ttsLanguages.
Example: ASR and Neural TTS only, with no transcription services
global:
asr:
enabled: true
transcribeRealtime:
enabled: false
transcribeBatch:
enabled: false
neuralTts:
enabled: true
Notes
- These settings sit directly under
global, alongsideenableNluandenableNeuron. global.enabledis still used only to turn whole charts on or off (lumenvoxSpeech,lumenvoxCommon,lumenvoxVb,externalServices).global.minimalInstall: truestill turns off ASR, transcription and TTS altogether, as in 7.x.
3. Neural TTS and legacy TTS
Neural TTS remains the default engine for every TTS language, as in 7.x. New in 8.0 is a single switch, global.neuralTts.enabled, that turns Neural TTS on or off for the whole deployment. Choosing the engine per language works as before, using legacyEnabled.
Voices must be listed explicitly in your values file so they are installed and available to the system.
Neural TTS (default) for a language
global:
neuralTts:
enabled: true # default
neuralttsDefaultVersion: "8"
ttsLanguages:
- name: "en_us"
voices:
- name: "aurora"
- name: "caspian"
Legacy TTS for a language
global:
ttsLanguages:
- name: "en_us"
legacyEnabled: true
voices:
- name: "chris"
Important: Turning Neural TTS off does not switch your languages to legacy TTS. With
neuralTts.enabled: false, a language only gets legacy TTS if it haslegacyEnabled: true. Any other language will have no TTS at all.
Notes
- Voice versions default to
neuralttsDefaultVersion("8") for neural voices andttsDefaultVersionfor legacy voices. You can set a specific version for an individual voice withversion. - The resource service now only advertises the voices for the engine that is actually deployed for each language.
- If you are upgrading Neural TTS from version 6.0.0, clear both the TTS cache folder and the Neural TTS models folder before starting the upgrade.
4. Real-time and batch transcription
Version 8.0.0 introduces two new services:
- transcribe-realtime replaces the previous streaming transcription components, with a re-architected, lower-latency pipeline.
- transcribe-batch is a new dedicated service for high-throughput offline transcription.
Both services are deployed for every ASR language by default. You can:
- Switch them off using the toggles in section 2.
- Turn on GPU acceleration for them in the same way as ASR (see section 1).
- Choose which models each one loads, per language (see section 5).
Optional real-time transcription tuning
lumenvox-speech:
transcribeRealtime:
enableFlowSense: false # set to true to enable FlowSense
europaBufferMs: "800" # audio buffer length, in milliseconds
Upgrading from 7.x: because both services are on by default, an upgraded installation gains these deployments (and their resource usage) for every ASR language unless you switch them off. Size your cluster accordingly.
Monitoring: transcription metric names now use the
transcribe_realtime_*ortranscribe_batch_*prefix. Update any monitoring dashboards and alerts.
5. ASR models per language
All ASR model settings for a language now live in that language's asrLanguages entry. ASR, real-time transcription and batch transcription each load only the models listed for them, rather than every model installed. This lets you run different models, or different model versions, on each service.
Example: English with a high-definition encoder and decoder, a fine-tuned model, and real-time transcription using the high-definition encoder only
global:
asrDefaultVersion: "8.0.0"
asrLanguages:
- name: "en"
version: "8.0.0"
fineTuned: true # renamed from enableFineTuned
services:
transcribeRealtime:
base: false # don't load the base encoder
extraModels:
- name: "asr_encoder_hidef_model_en"
version: "8.0.0"
# services omitted: loaded by asr, transcribeRealtime and transcribeBatch
- name: "asr_decoder_hidef_model_en_us"
version: "8.0.0"
services: ["asr"] # only asr loads this one
Settings
| Setting | Description |
|---|---|
extraModels | Adds models (such as high-definition variants) to a language. Use services on each model to limit which of asr, transcribeRealtime and transcribeBatch load it. If services is omitted, all three load it. |
services.<service>.base: false | Stops that service loading the language's base encoder. The base decoder is always loaded. |
fineTuned: true | Downloads and loads the language's fine-tuned model. This replaces enableFineTuned. |
version | Model version for a language or an individual model. If not set, asrDefaultVersion is used. |
asrDefaultVersion | Default version for any language or model without its own version. Now "8.0.0". |
global.customAsrModels | Now used only for shared download packages that aren't tied to a language. |
Safety check: the install fails with an error, rather than deploying a pod that repeatedly crashes, if any service ends up with no models for a language.
Upgrading from 7.x: this section contains three required changes. See items 2 to 4 in Upgrading from 7.x.
6. Istio Gateway API ingress
Version 8.0.0 adds a new ingress mode that uses the Kubernetes Gateway API with Istio. NGINX Ingress remains the default in the 8.0.0 charts and behaves exactly as in 7.x, including with custom Ingress class names such as nginx-public. No changes are needed to your NGINX settings if you continue using NGINX.
Switch to Istio Gateway API mode
global:
lumenvox:
ingress:
className: "istio" # any other value = Ingress with that class
gateway:
namespace: "istio-ingress"
createIstioNamespace: true # false if another release owns the namespace
disableTls: false # true when TLS is terminated at the load balancer
tlsSecretName: speech-tls-secret
infrastructureAnnotations: {} # cloud load balancer annotations
serviceExtraPorts: [] # e.g. port 443 for AWS NLB-terminated TLS
Requirements
- Istio installed (this provides the
istioGatewayClass). - Gateway API CRDs installed.
- Kubernetes 1.35 or later.
- Your TLS secret in the release namespace. The chart copies it into the gateway namespace automatically.
Configuration notes
| Topic | Details |
|---|---|
| Load balancer | Put your cloud load-balancer annotations in gateway.infrastructureAnnotations. Istio copies them onto the Service it creates. On Azure (provider: "azure"), the load-balancer health probe is configured automatically. |
| IP allowlist | In Istio mode, global.lumenvox.ingress.internalAllowlist (a list of CIDRs) restricts access to the admin portal, deployment portal, file-store, management API and reporting API. In NGINX mode, continue to restrict these routes as in 7.x, using a whitelist-source-range entry in httpAnnotations. |
| TLS | Istio mode uses gateway.disableTls. NGINX mode continues to use ingress.disableTls, as in 7.x. |
7. In-cluster databases for test and development
The new lumenvox-external-services subchart runs MongoDB, PostgreSQL, RabbitMQ and Redis inside your Kubernetes cluster. It replaces the Docker Compose setup previously used for test and development environments.
Not for production. The subchart's defaults are sized for a proof of concept. For production deployments, we continue to recommend managed or self-hosted database services. See the External Dependency Support Matrix for supported versions.
Step 1: Create the four secrets in the release namespace
kubectl create secret generic mongodb-existing-secret -n lumenvox --from-literal=mongodb-root-password=<password>
kubectl create secret generic postgres-existing-secret -n lumenvox --from-literal=postgresql-password=<password>
kubectl create secret generic rabbitmq-existing-secret -n lumenvox --from-literal=rabbitmq-password=<password>
kubectl create secret generic redis-existing-secret -n lumenvox --from-literal=redis-password=<password>
Passwords must be alphanumeric. The chart never creates these secrets for you; if one is missing, the install stops and shows the exact kubectl command needed to create it.
Step 2: Enable the subchart
global:
enabled:
externalServices: true
Notes
- Redis runs as a cluster by default, which requires the OT Redis Operator to be installed first. The install checks for the operator and shows the install command if it's missing. A single-instance mode is also available; see the subchart README.
- Storage class:set the storage class for your cluster:
lumenvox-external-services.storage.classNamefor MongoDB, PostgreSQL and RabbitMQ.lumenvox-external-services.redis.clusterMode.operator.storage.classNamefor the Redis cluster (for examplemanaged-csion AKS).- For clusters with no storage class at all, the subchart can install the local-path provisioner.
- Istio service mesh: if you use Istio as a service mesh, these database pods stay out of the mesh unless you set
global.serviceMesh.istio.injectDatabaseSidecars: true.
Upgrading from 7.x
Review these items before upgrading. Each one either stops the install with a message explaining what to change, or changes behaviour without warning.
Upgrade Kubernetes to 1.35 or later. Version 7.x required 1.33.
Rename
asrLanguages[].enableFineTunedtofineTuned. The old key stops the install with a message giving the new name.Move per-language models from
customAsrModelstoextraModels. In 7.x, a model listed incustomAsrModels(such as a high-definition encoder) was downloaded and loaded automatically. In 8.0, only models listed under a language are loaded, so per-language entries left incustomAsrModelsstop the install until they are moved under their language'sextraModels.Before (7.x):
global: customAsrModels: - name: "asr_encoder_hidef_model_en" asrLanguages: - name: "en" enableFineTuned: trueAfter (8.0):
global: asrLanguages: - name: "en" fineTuned: true extraModels: - name: "asr_encoder_hidef_model_en"Models without a version now download at 8.0.0.
asrDefaultVersionchanged from"7.0.0"to"8.0.0". To keep an older model, setversionexplicitly on the language or model.The default image tag is now
:8.0. Installations that don't setglobal.image.tagwill move to 8.0 images.Real-time and batch transcription are on by default. An upgraded installation gains these deployments for every ASR language unless you switch them off (see section 2).
Rollbacks: since version 7.x, rolling back to a pre-7.0 version with the standard helm rollback command is not supported. See How to Roll Back from Version 7.0 to 6.x.
If you have questions about your upgrade path, particularly if you are upgrading from a version earlier than 7.0, contact Capacity Support before proceeding.
