Skip to content

Deploy the GMRS ASR Pipeline on AirStack Edge

This tutorial deploys the reference GMRS ASR pipeline from the AirStack Model Vault. The decoupled Triton ensemble accepts GMRS narrowband-FM IQ, transcribes its audio with Voxtral, and conditionally emits compact transcript documents to an AirStack Edge webhook.

Pipeline

The GMRS ASR Pipeline ensemble has five models that are separately packaged and linked together with an ensemble model:

flowchart LR
  core[Radio\nDriver]
  demod[GMRS\nDemod]
  overlap[Audio\nOverlap]
  stt[Speech to\nText]
  stitcher[Transcript\nStitcher]
  batcher[Form\nMessage]
  endpoint[Webhook\nSink]

  subgraph edge [AirStack Edge]
    core -->|IQ| demod
    core -->|Metadata| demod
    subgraph triton [NVIDIA Triton Ensemble Model]
      demod -->|Audio| overlap
      demod -->|Metadata| overlap
      overlap -->|Audio| stt
      overlap -->|Metadata| stt
      stt -->|Text| stitcher
      stt -->|Metadata| stitcher
      stitcher -->|Text| batcher
      stitcher -->|Metadata| batcher
    end
  end
  batcher -->|JSON| endpoint
  1. fm_gmrs_demod uses GNU Radio blocks in a Triton Python model to demodulate interleaved float32 IQ into 16 kHz mono audio.
  2. audio_overlap retains two seconds of prior audio so transcription can span request boundaries.
  3. voxtral-mini-stt transcribes the overlapped audio.
  4. transcript_stitcher removes duplicate text due to audio overlap.
  5. fm_document_batcher accumulates new transcript text and emits compact document JSON with selected radio metadata.

The model names above are Triton logical names, not source-directory names. The ensemble is unbatched and decoupled. An accepted inference need not produce an OUTPUT response until the document batcher reaches its configured threshold.

Prerequisites

Before starting, make sure you have:

  • An AirStack Edge device with the inference server installed and running.
  • Access to the Client UI at https://<hostname>:<port>/ui/ and a bearer token.

  • Enough free storage for the six archives and the Voxtral model weights and packed environment.

AirStack Edge is the supported deployment path for these reference models. Follow the Triton inference application note for product-specific upload and model-loading details.

Step 1: Get the release source and Voxtral artifacts

Clone the repository revision you intend to deploy:

git clone https://github.com/deepwavedigital/airstack-edge-model-vault.git
cd airstack-edge-model-vault

The Voxtral model omits its weights and packed Python environment from Git. Download the two required files to their exact destinations:

wget -P models/speech_recognition/voxtral-mini-stt/1/ \
  https://archive.deepwavedigital.com/triton-models/voxtral-mini-stt/consolidated.safetensors
wget -P models/speech_recognition/voxtral-mini-stt/ \
  https://archive.deepwavedigital.com/triton-models/voxtral-mini-stt/voxtral-conda-env.tar.gz

Step 2: Package the six model archives

AirStack Edge accepts one archive per model. Each archive must be created at the content level: config.pbtxt, its numeric version directory, and all payload files are at the archive root, without a wrapping model-directory entry.

From the Model Vault root, create these six archives:

tar -czf fm_gmrs_demod.tar.gz -C models/demodulation/fm_gmrs_demod .
tar -czf audio_overlap.tar.gz -C models/stream_processing/audio_overlap .
tar -czf voxtral-mini-stt.tar.gz -C models/speech_recognition/voxtral-mini-stt .
tar -czf transcript_stitcher.tar.gz -C models/stream_processing/transcript_stitcher .
tar -czf fm_document_batcher.tar.gz -C models/stream_processing/fm_document_batcher .
tar -czf gmrs_asr_pipeline.tar.gz -C models/ensemble/gmrs_asr_pipeline .

For example, fm_gmrs_demod.tar.gz must contain config.pbtxt, 0/, and gmrs.py at its top level. Do not flatten the six models into one archive: the ensemble archive and every named dependency are uploaded and enabled separately.

Step 3: Upload and enable the models

Open the Client UI at https://<hostname>:<port>/ui/ and enter your bearer token.

  1. Open the Apps tab and select Upload.
  2. Upload each archive from Step 2: fm_gmrs_demod, audio_overlap, voxtral-mini-stt, transcript_stitcher, fm_document_batcher, and gmrs_asr_pipeline.
  3. Enable every model and wait for each status to show READY before starting the ensemble. Voxtral can take longer because Triton unpacks its custom environment while loading.

If a model does not reach READY, check the Console panel first. For lower-level errors, inspect the inference-server container:

docker logs airstack-edge-inference-server

Step 4: Test with a GMRS recording

The UI Verify controls inject a file directly into the selected model. Use a SigMF recording with GMRS/NBFM voice sampled at 1.92 MHz. The fm_gmrs_demod input is an even-length, interleaved real/imaginary float32 tensor; metadata is optional.

  1. On the Apps tab, select gmrs_asr_pipeline.
  2. Select the sigmf Verify control and choose a compatible SigMF data.
  3. Start verification and allow time for Voxtral inference and document batching.

The final stage is fm_document_batcher, not text_batcher. Its decoupled OUTPUT response is one JSON document when the buffered transcript reaches the deployed 400-character minimum or 2,000-character maximum. The model-level FLUSH control can emit partial text, but AirStack Edge does not supply that optional ensemble input in this workflow.

{
  "document": "Here's a sample transmission.",
  "doc_index": 0,
  "doc_metadata": {
    "streamId": "fm-462637500",
    "start_ntp_float": 1748448240.0,
    "end_ntp_float": 1748448246.0,
    "radio_metadata": {}
  }
}

No output before the threshold is expected; it is not an error.

Step 5: Run against a live receiver stream

After the recording test works, use the same ensemble with a live receiver stream.

  1. Set master_clock_rate to 61.44e6 and select Enable Device.
  2. Set data_type to c_f32, configure a 1.92e6 sample rate, and choose the GMRS channel frequency.
  3. Open the stream with Setup Stream and select gmrs_asr_pipeline as the model.
  4. Choose enough samples for several seconds of audio; six seconds is a reasonable starting point.
  5. Use continuous mode and configure at least one webhook for result delivery.
  6. Start the stream.

Continuous-mode document responses are delivered to the configured webhook. The final output is JSON; configure the webhook accordingly. AirStack Edge supplies the ensemble's IQ input and optional metadata in this workflow. It does not supply the optional RESET or FLUSH inputs, so document emission is threshold-based.

Extensions and model details

The five stages are independently reusable. You can replace the demodulator for another signal type or add a downstream document consumer while keeping the same AirStack Edge deployment pattern. Preserve the overlap/stitcher pair when using the checked-in GMRS pipeline: overlap gives Voxtral audio context and stitching removes repeated transcript text.

For exact model interfaces, configuration defaults, known limitations, artifact requirements, and deployment guidance, use the AirStack Model Vault and its GMRS pipeline guide{: target="_blank"}.