Release Notes 8.0.0
Release date: 25th September 2026
Previous release: Release Notes 7.1.0
Version 8.0.0 of Capacity Private Cloud is a major release that builds on the 7.1.0 baseline. It introduces a re-architected, lower-latency real-time transcription service with optional GPU acceleration, a new dedicated batch transcription service, a new Voice Activity Detection (VAD) engine, alongside a rebuilt English Inverse Text Normalization (ITN) engine, new ASR acoustic and language models, expanded SSML coverage, and a wide range of portal, grammar, TTS, licensing, and stability improvements.
This release affects Speech products. It is not available for Voice Biometric products; support for those will be made available in a future release.
New Features
The following capabilities were introduced in this release:
Real-time transcription service (
transcribe-realtime) — This new dedicated service includes a re-architected streaming transcription pipeline that reduces end-to-end latency, with selectable CPU or GPU execution. Note: these changes the deployment architecture and requires updates to your Helm values file - the following article provides information on where the GPU can be enabled at service and a language level Helm Chart Changes in Version 8.0.0. GPU execution additionally requires suitable NVIDIA GPU node resources. GPU processing provides latency improvements while allowing additional throughput. GPU mode is only currently available with high-definition acoustic models.Batch transcription service (
transcribe-batch) — A new dedicated service for high-throughput batch (offline) transcription, with optional GPU acceleration, VAD-based silence removal, and optimized audio ingestion (larger chunk sizes, no Redis streaming). The helm chart requirements on how to specify the different services and GPU enable is available here Helm Chart Changes in Version 8.0.0New Voice Activity Detection (VAD) engine — A new VAD model replaces the previous VAD, improving speech/silence detection accuracy, with improved handling of short utterances and initialization noise. VAD now contributes to dedicated VAD metrics:
snr_sensitivity(MRCP:snr-sensitivity-lvl),stream_init_delay(MRCP:vad-stream-init-delay) andvolume_sensitivity(MRCP:Sensitivity-Level). For ASR & Transcription services the VAD service is no longer required and has been built into the ASR, transcribed-realtime and transcribe-batch services. The VAD service is now only used for CPA & AMD interactions.New ASR/transcription acoustic and language models — Repackaged acoustic and language models enabled GPU processing for standard or high-definition models. Version 8.0 encoders and decoders are repackaged and versioned as 8.0 across the ASR family.
Rebuilt English ITN — English Inverse Text Normalization has been rebuilt on a unified text-normalizer module. Portuguese ITN now also uses this module.
Expanded SSML coverage — SSML support has been added for Dutch TTS voices and extended for Catalan/Valencian; and Italian currency.
Grammar pre-parse — Grammar loads now include a pre-parse step, improving validation and handling of grammars at load time.
Neuron event archiving — The archive service now captures Neuron (NLU) events, and can be retrieved via API
Separate Neural TTS RabbitMQ queue — Legacy and Neural TTS engines can now run for the same language, using a separate RabbitMQ queue for Neural TTS.
Azure AKS support for file-store — The file-store service now supports Azure AKS deployments.
Per-model ASR usage metrics — New Prometheus counters track ASR usage per model. Note: transcription metric names have changed to the
transcribe_realtime_*ortranscribe_batch_*prefix; update any monitoring dashboards accordingly.Istio API gateway support — The use of NGINX as the platform ingress provider has been deprecated and replaced with Istio API gateway. Additional information is available here:Helm Chart Changes in Version 8.0.0.
Neural TTS has been added as and optional install option within the helm charts.
Updates & Improvements
The following existing capabilities were improved or extended in this release:
Continuous transcription improvement - back-end improvements to continuous transcription were made that can result in improved accuracy.
Text normalization improvements — A broad set of English text-normalizer and ITN corrections, including punctuation and capitalization fixes, number/currency/address and long-number formatting, and resolution of normalization arithmetic errors.
Redaction improvements — Start/end time and duration are now reported for redacted words and abbreviation characters, and over-classification of ordinary numeric values as PII has been reduced.
Central model store & deterministic model loading — Models for ASR, transcription, ITN, and TTS are now loaded deterministically from a central model store per service. Note: this introduces mandatory per-service model-selection environment variables. Different models and version can now be applied to the various asr, transcribe-realtime or transcribe-batch services. Additional information is available here Helm Chart Changes in Version 8.0.0 to install the different models and different versions for asr, transcription and real-time.
Analysis Portal & Analysis Sets — Redesigned Analysis Sets review screens; TTS can now be selected as an analysis set option; interactions now display per-interaction timing and an audio scale; large analysis-set export & import has been improved.
Admin Portal version display — The Admin Portal now correctly displays the current deployment version, and the grammar and MRCP services now log their service version at startup.
Deployment Portal Diagnostics — The Diagnostics page has been updated to include tests for Neural TTS, real-time transcriptions, batch transcriptions, CPA and AMD.
Deployment Portal TTS tester — The TTS tester now allows for tests to be conducted using legacy and neural TTS if loaded in the same environment.
Licensing / metrics — Updated port-level license statistics, total session request counters, and other licensing/metrics reporting inaccuracies.
Diarization: Various improvements have been made to improve the overall performance & accuracy
Go SDK session events — The Go SDK now returns additional session events, including a clear "service not available" event when a dependent service is unavailable.
Bug Fixes
The following defects were resolved in this release:
Legacy (CPU) TTS — Fixed the Legacy TTS service failing to start (missing
libpcrecpp.so.0and internal synth binary naming/connection issues) which caused all synthesis to fail.Neural TTS — Fixed a memory leak in the session service under Neural TTS load; fixed an SSML
<audio>URI being replayed once per nested<voice>element; fixed missing word/sentence offsets for long SSML text; fixed intermittent missing voice offsets in multi-voice synthesis; fixed a crash on SSML<phoneme>supplied without aphattribute.TTS SSML address reading — Fixed
say-as interpret-as="address"speaking a street number as a whole number instead of individual digits; corrected large SSML prefetch handling.ASR recognition accuracy — Fixed streaming recognition regressions affecting English (e.g. "apple pay" → "applepay"), Spanish, and word-onset timing, and restored recognition accuracy affected by an engine/decoder change.
Grammar under load — Fixed invalid GRXML being reported as loaded and a subsequent load hanging a worker; and reduced excessive memory use on very large grammar loads.
Grammar over HTTPS — Fixed grammar preload TLS fetch failures nvs89ot being surfaced (hang / no-input), and
ssl_verify_peer=falsenot being honored.Neuron (NLU) — Fixed slots not being populated and added a clear error when an invalid or uninstalled model is referenced (previously the request could hang).
MRCP — Fixed a container crash on startup (missing
libpcrecpp.so.0); fixed a CPA RECOGNIZE that never terminated when the audio stream stopped before classification; fixed ASR no-match on digit recognition; fixed an inline transcription grammar being accepted but never resolved; corrected NLSML result differences observed between v6.3 and v7.1. Note: integrations that parse NLSML output should validate against this release.ITN digit formatting — Fixed text normalization no longer formatting phone/SSN digits into their expected dashed format.
Deployment Portal — Fixed missing-partition-table errors when viewing ASR interactions with notes; fixed non-English voices not appearing in the TTS Tester for some locales.
Reporting API — Fixed a License Reporter Tool failure and a case where a batched license usage report was dropped when a batch failed.
ASR reliability — Fixed
bargeinThresholdvalues of 90+ silently failing with a misleading "timed out in audio puller" error; and an intermittent batch recognition request loss where the engine never receivedStartAsrInteraction.Built-in grammars — Fixed an en-US built-in time bug affecting the
sloppy_minusrule and corrected incorrect values returned by the built-in en-AU currency grammar.Continuous / real-time transcription — Resolved several streaming/continuous transcription defects: VAD-active streaming never finalizing (error code 13); partial results not delivered; barge events delivered with a rewound audio offset; and a continuous audio-starvation terminal error not surfaced to the SDK/client
Call-flow timing — Fixed
PROMPT_ENDbeing detected roughly 9 seconds after the end of the greeting in an Apple Call Screen scenario.Storage service logging — Fixed info-level logs being emitted at startup regardless of the configured log level.
Security — Resolved Snyk-detected vulnerabilities across various images
Patch Releases
The following patch releases were issued against version 7.1.0 containers and are included in the 8.0.0 baseline. Customers running 7.1.x patch releases do not need to apply these separately, and are instead encouraged to upgrade to this version, which includes additional improvements.
Service | Patch Version. | Description |
|---|---|---|
asr | 7.1.1 | ASR timestamp accuracy (half-scale word timestamps / subsampling factor); enhanced-transcription regression; EN recognition regressions from mid-cycle engine builds; recognition-accuracy collapse from a decoder hotfix that dropped tokens; 7.1 rollback regressions; excessive memory (~12 GB) on large grammar loads; release image validation & packaging. |
grammar | 7.1.1 | Resolved the production grammar service losing RabbitMQ connectivity and ceasing to work. |
grammar | 7.1.2 | Resolved grammar pod restarts observed in production. |
grammar | 7.1.3 | Malformed grammar that could crash the grammar manager; incorrect built-in en-AU currency grammar values; unexpected MATCH/NO-MATCH via MRCP-API. |
grammar | 7.1.4 | Grammar load error; grammar pod SSL certificate verification failures after upgrading to 7.1. |
vad | 7.1.2 | CPA screening-tone classification returned "UNKNOWN_SILENCE" for LINEAR16 audio (correct "UNKNOWN_SPEECH" for ULAW). |
vad | 7.1.3 | PROMPT_END detected ~9 seconds after the end of the greeting (Apple Call Screening). |
vad | 7.1.4 | Fixed prompt-end delay for MRCP interactions. |
neural-tts | 7.1.1 | Fallback text / pre-recorded audio incorrectly played; added measures support to the Basque normaliser. |
neural-tts | 7.1.2 | Issue when performing a large SSML prefetch. |
neural-tts | 7.1.3 | Neural TTS replayed an SSML <audio> URI once per nested <voice> element. |
session | 7.1.1 | Inline TTS requests that failed to reach the Neural TTS service. |
session | 7.1.4 | Continuous-transcription returned Error Code 13 for almost every interaction on the latest 7.1 ASR patch. |
mrcp | 7.1.1 | Incorrect result when an enhanced-transcription grammar and a DTMF grammar were supplied together with DTMF. |
mrcp-api | 7.1.1 | NLSML results changed between v6.3 and v7.1 (and ported to 7.1 production). |
license | 7.1.1 | License Reporter Tool failure in the Reporting API. |
deployment | 7.1.1 | Handling of PostgreSQL database passwords containing special characters. |
deployment-portal | 7.1.1 | Non-English voices not present in the TTS Tester dropdown for some locales. |
The coordinated 7.1.1 and 7.1.2 patch releases also bundled additional ITN / text-normalizer corrections (compound-number, "trillion dollar" decimal formatting and arithmetic errors), a fix for a large grammar that could cause ASR service failure, and VAD barge-in / session-close handling for a customer use case.
Installation Notes
Before installing or upgrading to version 8.0.0, review the following requirements. For full step-by-step installation guidance, refer to the Setting Up a Deployment article and the relevant setup guides in the Installation section of this knowledge base.
Infrastructure Requirements
The following are the recommended dependencies when using this version of the product.
kubeadm v1.35
MongoDB v8.3.4
Postgres v17.9
Redis v8.6.6
RabbitMQ v4.2.3
And operating systems:
Almalinux 8 and 9
Debian 12
RHEL 8 and 9
Rocky 8 and 9
Ubuntu 22.04, 24.04 and 26.04
The use of NGINX as the platform ingress provider has been deprecated and replaced with Istio API gateway. Additional information is available here: Helm Chart Changes in Version 8.0.0
See our External Dependency Support Matrix article for more information.
Concurrency Limits
To maintain stable performance at scale, customers are encouraged to utilize the following per-pod concurrency limits, however individual use-cases vary, so adjustments for each specific use-case may be necessary:
session & lumenvox-api pods: maximum 400 concurrent interactions per pod
vad: maximum 300 concurrent interactions per pod
grammar pods: maximum 100 concurrent interactions per pod
itn : maximum 48 concurrent interactions per pod
Usage Recommendations
TTS text input: recommended maximum 4 MB per request
Transcription audio: a maximum of 120 minutes per session is supported
Helm Chart Setup
Run the following command to update your local Helm chart repository before installation:
See our public Helm Charts repository for the latest and previously released versions, along with more information.helm repo update
For MRCP deployments: there is no Helm chart for MRCP. Use the provided Docker Compose file available in the mrcp-api repository. The MRCP service runs on its own Docker virtual machine which then communicates with the configured Kubernetes cluster.
Real-Time & Batch Transcription — New in 8.0.0
Version 8.0.0 introduces the new transcribe-realtime service (replacing the previous streaming transcription components) and a new transcribe-batch service for offline/high-throughput transcription. Both must be enabled in your Helm values file. Execution can run on CPU or GPU, selectable per deployment; GPU mode requires the GPU-enabled distribution packages and appropriate node resources. Review the updated architecture and values file before installing or upgrading.
Model Selection Environment Variables — New in 8.0.0
Models are now loaded deterministically from a central model store per service. Version 8.0 introduces mandatory per-service model-selection environment variables for encoder/decoder selection (for example ASR_ENCODERS/ASR_DECODERS, TRANS_RT_ENCODERS/TRANS_RT_DECODERS, TRANS_BATCH_ENCODERS/TRANS_BATCH_DECODERS). Ensure these are set for each relevant service in your values file. Confirm the exact variable names and values against the 8.0 deployment guide.
Neural TTS Configuration
To configure Neural TTS in your values file, specify the legacyEnabled flag per language. Set to false to enable Neural TTS (recommended); set to true to use legacy TTS. Neural TTS voice models must be explicitly listed in the values file so they can be installed and available to the system. Legacy and Neural TTS engines can now run for the same language. TTS will route to the specific engine based on the voice selected, alternatively the engine can be specified when creating a TTS interaction.
ttsLanguages: - name: "en_us" legacyEnabled: false voices: - name: "jeff"
ITN Configuration
When deploying ITN, include the relevant language entries in your Helm values file (ensure ITN languages are versioned):
itnLanguages: - name: "en" - name: "es"
pt for Portuguese, de for German) as required by your deployment.Upgrade Procedures
Upgrading from previous versions of Capacity Private Cloud is supported. Contact Capacity Support to discuss your specific upgrade path before proceeding, particularly if upgrading from a version earlier than 7.0. Further information can be found here: http://privatecloud.capacity.com/article/733447
Architecture change: Because 8.0.0 introduces the new transcribe-realtime and transcribe-batch services, an updated VAD, and mandatory model-selection environment variables, review the updated Helm values and architecture before upgrading. Deployments using GPU acceleration require the GPU-enabled distribution packages and suitable node resources.
If upgrading Neural TTS from version 6.0.0, the TTS cache folder and the Neural TTS models folder must both be cleared before starting the upgrade. Contact support if you have questions about this step.
API Reference
See our API documentation for 8.0.0 for details about this version's API.
Model Versions
Various acoustic, language and other AI models are distributed for use with our services, and the versioning of those models does not follow the same version as the corresponding software release, since the models may be modified less frequently. The following model versions were introduced and designed for this release:
ASR — 8.x: All existing standard and high-definition models encoder & decoder models repackaged to enable GPU processing.
ITN models updated: English ITN rebuilt (updated dataset); Portuguese ITN migrated to the unified text-normalizer module
