Samples

Two samples are provided in the Cochl.Sense Edge SDK: sense-file and sense-stream.

  • sense-file runs inference on an audio file (WAV or MP3).
  • sense-stream runs inference on a live audio buffer (e.g. a microphone).

Make sure your environment is set up first — see Getting Started. Every platform below shares the same three preparation steps — clone the tutorial, download the SDK, unzip the SDK into the tutorial — and needs these command-line tools. On Linux hosts:

sudo apt update
sudo apt install -y git unzip curl

Then pick the platform you’re targeting. Each tab follows the same order: Requirements → Prepare → License manager → Usage → Build & run → Reference → Removed in v1.6.1 → Troubleshooting. Android has no License manager step (the license runs in-process) and no Troubleshooting section.

1. Requirements

The Cochl.Sense Edge SDK for Python supports Python 3.6–3.12. The wheel is a compiled, version-specific build, so you must download the file that matches both your Python version and CPU architecture (e.g. cp38 for Python 3.8, aarch64 for a 64-bit Raspberry Pi). If you need a version that is not listed, contact support@cochl.ai.

In addition to the shared git unzip curl tools above, install the Python build tools and the PortAudio library the sense-stream sample uses for microphone capture:

sudo apt install -y python3-venv python3-pip libportaudio2

Check your interpreter version first — you’ll download the matching SDK in step 2:

python3 --version   # e.g. Python 3.8.10  -> download the cp38 wheel

2. Prepare the sample

(1) Clone the tutorial

git clone https://github.com/cochlearai/sense-sdk-python-tutorials.git

(2) Create and activate a virtual environment

Create the environment with the interpreter whose version matches the SDK you’re about to download:

cd sense-sdk-python-tutorials
python3 -m venv ./venv
source ./venv/bin/activate     # prompt is now prefixed with (venv)
python -m pip install --upgrade pip
  • NOTE: For older Python versions (e.g. 3.6, 3.7) the venv module ships in a version-specific package (python3.6-venv, python3.7-venv). Install the one matching your interpreter.

(3) Download the version-matched SDK

Download the wheel package built for your Python version and architecture. Replace cp38 and linux_aarch64 with the values matching your environment — note Python 3.6 and 3.7 use the ABI tags cp36-cp36m / cp37-cp37m, not cp36-cp36 / cp37-cp37. Resources lists the exact link for every version:

# (venv)
curl -LO https://github.com/cochlearai/sense-sdk-python-tutorials/releases/download/v1.6.1/sense-1.6.1-cp38-cp38-linux_aarch64.zip

(4) Unzip the SDK into the tutorial directory

unzip sense-1.6.1-cp38-cp38-linux_aarch64.zip -d .

After unzipping, your sample directory should look like this:

sense-sdk-python-tutorials
    |-- README.md
    |-- config.json          # configuration file used to initialize the SDK
    |-- audio-files          # audio samples
    |-- sense-file           # sense-file sample (sense_file.py)
    |-- sense-stream         # sense-stream sample (sense_stream.py)
    `-- sense                # Cochl.Sense Edge SDK
        |-- sense-1.6.1-cp38-cp38-linux_aarch64.whl
        `-- sl               # license daemon (scripts + binary)

(5) Install the SDK and dependencies

Install the wheel you unzipped. The Sense wheel has no runtime dependencies; the sense-stream sample additionally needs sounddevice (from its requirements.txt):

# (venv)
pip install sense/sense-1.6.1-cp38-cp38-linux_aarch64.whl
pip install -r sense-stream/requirements.txt   # sense-stream only (sounddevice)

3. Install the license manager

The license manager is required from SDK v1.4.0. From v1.6.1 it is the native sense-lm backend, which replaced Thales. On Linux it runs as a small system daemon that verifies the license and needs an internet connection. Install it once:

sudo ./sense/sl/scripts/install_daemon.sh

The installer creates a system group sense-lm and adds your user to it. Apply the new group membership without logging out:

newgrp sense-lm     # or log out and back in

To remove the daemon later:

sudo ./sense/sl/scripts/uninstall_daemon.sh

4. How to use the Cochl.Sense Edge SDK for Python

(1) Initialization

Initialize the SDK with your project key and the contents of config.json. sense.session(...) is a context manager that calls init() on entry and terminate() on exit. Initialization may take some time if a model must be downloaded (the SDK downloads a model when none is present or a newer one is available).

import sense

project_key = "Your project key"  # Replace with your own key
with open("config.json", encoding="utf-8") as f:
    config = f.read()

with sense.session(project_key, config):
    # ... create a processor and run inference here ...
    pass

(2) Configuration

This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.

{
  "device_name": "SenseDevice",
  "num_threads": 0,
  "log_level": "info",
  "app_data_path": "~",
  "default_hopsize": 1.0,
  "metrics": {
    "retention_period": 0,
    "free_disk_space": 100,
    "push_period": 30
  },
  "sensitivity_control": {
    "default_sensitivity": "NORMAL",
    "tag_sensitivity": {
    }
  },
  "result_summary": {
    "enable": false,
    "default_interval_margin": 0,
    "tag_interval_margin": {
    }
  },
  "audio_preprocessor": {
    "audio_activity_detection": {
      "enable": false,
      "history_count": 600,
      "sensitivity": "NORMAL"
    },
    "automatic_gain_control": {
      "enable": false
    }
  }
}

Top-level fields

FieldDefaultDescription
device_namerequiredName shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page.
num_threads0Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1).
log_level"info"Log verbosity, as a string: "error", "warn", "info", "debug".
app_data_path"~"Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it.
default_hopsize1.0Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0].
metricsobjectLocal metrics buffer + push behavior. See below.
sensitivity_controlobjectDetection sensitivity, globally and per tag. See below and Advanced Configurations.
result_summaryobjectMerge consecutive same-tag windows into single interval events. See below and Advanced Configurations.
audio_preprocessorobjectAudio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations.

metrics sub-fields

Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.

FieldDefaultUnitDescription
retention_period0daysHow long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days.
free_disk_space100MBWhen free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB.
push_period30secondsHow often metrics are pushed to the dashboard. Capped at 3,600 s.

sensitivity_control sub-fields

Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".

FieldDefaultDescription
default_sensitivity"NORMAL"Detection sensitivity applied to every tag.
tag_sensitivity{}Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity.

result_summary sub-fields

When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).

FieldDefaultUnitDescription
enablefalse—Turns Result Summary on.
default_interval_margin0secondsGap tolerance for merging same-tag windows into one interval.
tag_interval_margin{}secondsPer-tag margin overrides, e.g. { "Speech": 2 }.

audio_preprocessor sub-fields

Both preprocessors apply to stream mode only.

FieldDefaultDescription
audio_activity_detection.enablefalseSkips inference on windows with no audio activity.
audio_activity_detection.history_count600Detection window length, in seconds (a fixed span, independent of hop size).
audio_activity_detection.sensitivity"NORMAL"Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH".
automatic_gain_control.enablefalseNormalizes input loudness before inference.

(3) Audio input and predict

The SDK receives audio and returns detected sound tags. In file mode you point it at a file; in stream mode you push captured audio buffers. Results are delivered to a sense.ResultListener callback (on_result), one call per inference window.

Sample rate. The input must be at or above the model’s rate (typically 22,050 Hz). Lower rates are not supported — file mode rejects a below-rate file with a resampling error and stream mode rejects the pushed chunk; higher rates are automatically downsampled. Use processor.model_sample_rate() to capture at the model’s native rate.

Create a file processor and start it. start() blocks until the whole file is processed.

Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.

import sense

class Printer(sense.ResultListener):
    def on_result(self, frame: sense.Result) -> None:
        if not frame.is_ok():
            print("error:", frame.error)
            return
        for tag in frame.tags:
            print(frame.start_time, tag.name, tag.probability)

listener = Printer()
processor = sense.create_file_processor("audio_file.wav", listener)
processor.start()   # blocks until end of file

Create a stream processor, then push captured audio chunks. Use the processor as a context manager (with processor:) to arm/stop the stream.

Note: connect a microphone (or other input device) before running the stream sample. This tutorial uses sounddevice (PortAudio) for capture; you can replace it with any source.

import sounddevice as sd
import sense

listener = Printer()                                   # sense.ResultListener
processor = sense.create_stream_processor(listener)
rate = processor.model_sample_rate()                   # capture at the model rate

mic = sd.RawInputStream(samplerate=rate, channels=1, dtype="int16")
with mic, processor:                                   # processor.start() arms the stream
    while True:
        data, _ = mic.read(rate)                       # one hop of audio
        processor.push(bytes(data), 1, sense.SampleFormat.INT16, rate)

(4) Result format

Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.

Your ResultListener.on_result is called once per window with a sense.Result:

class Tag:
    name: str                  # Detected tag, e.g. "Siren"
    probability: float         # Confidence for that tag, 0.0-1.0

class Result:
    tags: List[Tag]            # Empty when nothing passed threshold
    start_time: float          # Window start, seconds
    end_time: float            # Window end, seconds
    prediction_time_ms: float  # Processing time for this window, ms
    timestamp: int             # Unix time (seconds) the window was processed
    summaries: List[str]       # Result Summary lines; empty unless enabled
    error: str                 # Empty on success, message on failure

    def is_ok(self) -> bool: ...   # True when error is empty
FieldTypeNotes
tagsList[Tag]Detected tags; empty for silence or sub-threshold audio — a valid, successful result.
start_timefloatSeconds from the start of the stream or file.
end_timefloatstart_time plus the window length.
prediction_time_msfloatInference latency for this window.
timestampintUnix time, in seconds, when the window was processed.
summariesList[str]Interval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval.
errorstrEmpty on success; non-empty only on failure, with tags then empty.

Check result.is_ok() rather than testing tags — an empty tags means silence, not an error.

(5) Terminate

The SDK allocates resources during initialization. Release them with terminate(). When you use sense.session(...), this happens automatically on exit; call it explicitly only if you used sense.init(...) directly.

import sense

sense.terminate()

5. How to run

Run from the tutorial root so the samples find config.json in the working directory. Set your project key in PROJECT_KEY at the top of sense-file/sense_file.py and sense-stream/sense_stream.py first.

# (venv)
python sense-file/sense_file.py <PATH_TO_AUDIO_FILE>
python sense-file/sense_file.py audio-files/babycry.wav

Note: connect a microphone before running.

# (venv)
python sense-stream/sense_stream.py     # Ctrl-C to stop

6. Reference

Import from the sense package.

Module functions

sense.init()
sense.init(project_key, config_json)

Authenticates with the project key and applies the configuration.

ParameterTypeDescription
project_keystrYour project key.
config_jsonstrThe contents of config.json (not a path).

Returns bool. Raises RuntimeError on failure — it does not return False, so if not sense.init(...) never fires. Wrap calls in try / except RuntimeError, or use with sense.session(...), which does it for you.

sense.terminate()
sense.terminate()

Releases all resources allocated during initialization. Returns bool — True on success.

sense.is_initialized()
sense.is_initialized()

Returns bool — whether the SDK is initialized.

sense.session()
sense.session(project_key, config_json)

Context manager that calls init() on entry and terminate() on exit: with sense.session(key, config): .... Parameters are the same as init. Returns a context manager.

sense.create_file_processor()
sense.create_file_processor(file_path, listener)

Creates a file-mode processor.

ParameterTypeDescription
file_pathstrPath to the audio file (WAV or MP3).
listenerResultListenerReceives one on_result call per window.

Returns Processor. Raises RuntimeError on failure.

sense.create_stream_processor()
sense.create_stream_processor(listener)

Creates a stream-mode processor fed by push.

ParameterTypeDescription
listenerResultListenerReceives one on_result call per window.

Returns Processor. Raises RuntimeError on failure.

Processor

Processor.start()
Processor.start()

File mode runs to completion (blocks); stream mode arms the stream. Returns bool. Raises RuntimeError on failure.

Processor.push()
Processor.push(pcm, num_channels, sample_format, input_sample_rate)

Feeds one chunk of captured audio (stream mode); the SDK buffers it internally and runs inference only once a full model window has accumulated — so a single push may deliver zero, one, or several results.

ParameterTypeDescription
pcmbytesInterleaved PCM for one hop.
num_channelsintChannel count (≥ 1).
sample_formatintA SampleFormat value.
input_sample_rateintCapture rate in Hz; must be ≥ model_sample_rate().

Returns bool — True if the chunk was accepted.

Processor.stop()
Processor.stop()

Stops processing. Returns bool.

Processor.model_sample_rate()
Processor.model_sample_rate()

Returns int — the model’s native sample rate in Hz.

Runtime controls

Runtime feature controls on the processor. AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".

Setter / GetterTypeDescription
set_sensitivity(level) / get_sensitivity()strGlobal detection sensitivity.
set_tag_sensitivity(tag, level) / get_tag_sensitivity(tag)strPer-tag sensitivity.
set_result_summary_enabled(enable) / is_result_summary_enabled()boolMerge same-tag windows into intervals.
set_result_summary_margin(seconds) / get_result_summary_margin()intInterval merge margin (seconds).
enable_audio_activity_detection(enable) / is_audio_activity_detection_enabled()boolSkip inference on inactive audio.
set_aad_sensitivity(level) / get_aad_sensitivity()strActivity-detection sensitivity.
set_aad_history_count(seconds) / get_aad_history_count()intActivity window length (seconds).
enable_automatic_gain_control(enable) / is_automatic_gain_control_enabled()boolNormalize input loudness.

ResultListener

Subclass and override on_result(self, frame: Result) -> None. It fires once per inference window on the SDK’s inference thread, so keep the callback cheap.

Result

Passed to on_result. is_ok() -> bool reports whether the window succeeded.

FieldTypeDescription
tagsList[Tag]Detected tags; each Tag has name: str, probability: float.
start_time / end_timefloatWindow bounds, in seconds.
prediction_time_msfloatProcessing time, in milliseconds.
timestampintWindow timestamp.
summariesList[str]Interval lines (present when Result Summary is enabled).
errorstrError message (empty when is_ok()).

SampleFormat

SampleFormat.FLOAT32 | INT16 | INT32 | FLOAT64 — the encoding passed to push. FLOAT64 is for file/pre-recorded input only.

7. Removed in v1.6.1

The pre-1.6 procedural API was replaced by the processor/listener API above. Migrate as follows:

RemovedUse instead
SenseInit(project_key, config_path)sense.init(project_key, config_contents) or sense.session(...)
SenseTerminate()sense.terminate()
AudioSourceFile().Load() + .Predict()sense.create_file_processor(path, listener) + processor.start()
AudioSourceStream().Predict(buffer, rate)sense.create_stream_processor(listener) + processor.push(...)

8. Troubleshooting

If you hit an error, reset the SDK’s local state and run again:

sudo rm -rf ~/.sense-sdk

1. Requirements

In addition to the shared git unzip curl tools above, the C++ sample needs a C++14 compiler and CMake (≥ 3.13, for the cmake -B option used below):

sudo apt install -y build-essential cmake

2. Prepare the sample

(1) Clone the tutorial

git clone https://github.com/cochlearai/sense-sdk-cpp-tutorials.git

(2) Download the SDK

Download the Linux build for your architecture (replace x86_64 with aarch64 on 64-bit Arm):

cd sense-sdk-cpp-tutorials
curl -LO https://github.com/cochlearai/sense-sdk-cpp-tutorials/releases/download/v1.6.1/sense-sdk-1.6.1-linux-x86_64-cpp.zip

(3) Unzip the SDK into the tutorial directory

unzip sense-sdk-1.6.1-linux-x86_64-cpp.zip -d .

After unzipping, your sample directory should look like this:

sense-sdk-cpp-tutorials
    |-- CMakeLists.txt
    |-- README.md
    |-- config.json          # configuration file used to initialize the SDK
    |-- audio-files          # audio samples
    |-- sense-file           # sense-file sample (sense_file.cc)
    |-- sense-stream         # sense-stream sample (sense_stream.cc)
    `-- sense                # Cochl.Sense Edge SDK
        |-- include
        |-- lib
        `-- sl               # license daemon (scripts + binary)

3. Install the license manager

The license manager is required from SDK v1.4.0. From v1.6.1 it is the native sense-lm backend, which replaced Thales. On Linux it runs as a small system daemon that verifies the license and needs an internet connection. Install it once:

sudo ./sense/sl/scripts/install_daemon.sh

The installer creates a system group sense-lm and adds your user to it. Apply the new group membership without logging out:

newgrp sense-lm     # or log out and back in

To remove the daemon later:

sudo ./sense/sl/scripts/uninstall_daemon.sh

4. How to use the Cochl.Sense Edge SDK for C++

(1) Initialization

Initialize the SDK with your project key and the contents of config.json (read the file yourself and pass the string). sense::Init returns false on failure and writes a message to the optional error pointer. Initialization may take some time if a model must be downloaded.

#include <string>
#include "sense/sense.hpp"

std::string project_key = "Your project key";  // Replace with your own key
std::string config = /* contents of config.json */;

std::string err;
if (!sense::Init(project_key, config, &err)) {
    // Init has failed; see err.
    return -1;
}

(2) Configuration

This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.

{
  "device_name": "SenseDevice",
  "num_threads": 0,
  "log_level": "info",
  "app_data_path": "~",
  "default_hopsize": 1.0,
  "metrics": {
    "retention_period": 0,
    "free_disk_space": 100,
    "push_period": 30
  },
  "sensitivity_control": {
    "default_sensitivity": "NORMAL",
    "tag_sensitivity": {
    }
  },
  "result_summary": {
    "enable": false,
    "default_interval_margin": 0,
    "tag_interval_margin": {
    }
  },
  "audio_preprocessor": {
    "audio_activity_detection": {
      "enable": false,
      "history_count": 600,
      "sensitivity": "NORMAL"
    },
    "automatic_gain_control": {
      "enable": false
    }
  }
}

Top-level fields

FieldDefaultDescription
device_namerequiredName shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page.
num_threads0Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1).
log_level"info"Log verbosity, as a string: "error", "warn", "info", "debug".
app_data_path"~"Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it.
default_hopsize1.0Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0].
metricsobjectLocal metrics buffer + push behavior. See below.
sensitivity_controlobjectDetection sensitivity, globally and per tag. See below and Advanced Configurations.
result_summaryobjectMerge consecutive same-tag windows into single interval events. See below and Advanced Configurations.
audio_preprocessorobjectAudio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations.

metrics sub-fields

Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.

FieldDefaultUnitDescription
retention_period0daysHow long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days.
free_disk_space100MBWhen free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB.
push_period30secondsHow often metrics are pushed to the dashboard. Capped at 3,600 s.

sensitivity_control sub-fields

Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".

FieldDefaultDescription
default_sensitivity"NORMAL"Detection sensitivity applied to every tag.
tag_sensitivity{}Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity.

result_summary sub-fields

When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).

FieldDefaultUnitDescription
enablefalse—Turns Result Summary on.
default_interval_margin0secondsGap tolerance for merging same-tag windows into one interval.
tag_interval_margin{}secondsPer-tag margin overrides, e.g. { "Speech": 2 }.

audio_preprocessor sub-fields

Both preprocessors apply to stream mode only.

FieldDefaultDescription
audio_activity_detection.enablefalseSkips inference on windows with no audio activity.
audio_activity_detection.history_count600Detection window length, in seconds (a fixed span, independent of hop size).
audio_activity_detection.sensitivity"NORMAL"Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH".
automatic_gain_control.enablefalseNormalizes input loudness before inference.

(3) Audio input and predict

Build an AudioProcessor with a source type and a FrameResultListener. Each inference window is delivered to your listener’s OnResult callback.

Sample rate. The input must be at or above the model’s rate (typically 22,050 Hz). Lower rates are not supported — file mode rejects a below-rate file with a resampling error and stream mode rejects the pushed chunk; higher rates are automatically downsampled. Use processor->model_sample_rate() to capture at the model’s native rate.
#include "sense/sense.hpp"

class Printer : public sense::FrameResultListener {
 public:
  void OnResult(const sense::FrameResult& result) override {
    if (!sense::IsOk(result)) { /* handle result.error */ return; }
    for (const auto& tag : result.tags) { /* tag.name, tag.probability */ }
  }
};

Build a kFile processor and start it. StartProcessing() blocks until the whole file is processed.

Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.

#include <memory>
#include "sense/sense.hpp"

Printer listener;
std::string build_err;
std::unique_ptr<sense::AudioProcessor> processor =
    sense::AudioProcessorBuilder()
        .SetSourceType(sense::AudioSourceType::kFile)
        .SetFilePath("audio_file.wav")
        .SetCallback(&listener)
        .Build(&build_err);
if (!processor) { /* handle build_err */ }

std::string start_err;
processor->StartProcessing(&start_err);   // blocks until end of file

Build a kRecorder processor, arm it with StartProcessing(), then push captured audio via PushAudioChunk().

Note: connect a microphone (or other input device) before running the stream sample. The tutorial captures audio itself (vendored miniaudio); the SDK has no audio-device dependency.

#include <memory>
#include "sense/sense.hpp"

Printer listener;
std::unique_ptr<sense::AudioProcessor> processor =
    sense::AudioProcessorBuilder()
        .SetSourceType(sense::AudioSourceType::kRecorder)
        .SetCallback(&listener)
        .Build();
processor->StartProcessing();               // arms the stream
const int rate = processor->model_sample_rate();

// For each captured chunk (interleaved PCM), push it to the SDK:
processor->PushAudioChunk(
    reinterpret_cast<const uint8_t*>(buffer), size_in_bytes,
    /*num_channels=*/1, sense::SampleFormat::kFloat32, rate);

(4) Result format

Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.

Your FrameResultListener::OnResult is called once per window with a sense::FrameResult:

struct FrameResult {
  struct Tag {
    std::string name;            // Detected tag, e.g. "Siren"
    float       probability;     // Confidence for that tag, 0.0-1.0
  };

  std::vector<Tag>         tags;                // Empty when nothing passed threshold
  float                    start_time;          // Window start, seconds
  float                    end_time;            // Window end, seconds
  double                   prediction_time_ms;  // Processing time for this window, ms
  int64_t                  timestamp;           // Unix time (seconds) the window was processed
  std::vector<std::string> summaries;           // Result Summary lines; empty unless enabled
  std::string              error;               // Empty on success, message on failure
};
FieldTypeNotes
tagsstd::vector<Tag>Detected tags; empty for silence or sub-threshold audio — a valid, successful result.
start_timefloatSeconds from the start of the stream or file.
end_timefloatstart_time plus the window length.
prediction_time_msdoubleInference latency for this window.
timestampint64_tUnix time, in seconds, when the window was processed.
summariesstd::vector<std::string>Interval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval.
errorstd::stringEmpty on success; non-empty only on failure, with tags then empty.

Check IsOk(result) rather than testing tags — an empty tags means silence, not an error.

(5) Terminate

The SDK allocates resources during initialization. Release them with sense::Terminate().

#include "sense/sense.hpp"

sense::Terminate();

5. Build and run

The tutorial ships a CMakeLists.txt. Set your project key in kProjectKey at the top of sense-file/sense_file.cc and sense-stream/sense_stream.cc, then build:

cmake -B build
cmake --build build

This produces sense-file-app and sense-stream-app next to config.json and audio-files/, so you can run them in place.

./sense-file-app <PATH_TO_AUDIO_FILE>
./sense-file-app audio-files/babycry.wav

Note: connect a microphone before running.

./sense-stream-app       # Ctrl-C to stop

6. Reference

Declared in sense/sense.hpp (namespace sense).

Free functions

Init()
bool Init(const std::string& project_key, std::string config_json, std::string* error = nullptr)

Authenticates with the project key and applies the configuration.

ParameterTypeDescription
project_keyconst std::string&Your project key.
config_jsonstd::stringThe contents of config.json (not a path).
errorstd::string*Optional; receives a message on failure.

Returns bool — true on success.

Terminate()
bool Terminate() noexcept

Releases all resources allocated during initialization. Returns bool.

IsInitialized()
bool IsInitialized() noexcept

Returns bool — whether the SDK is initialized.

sense::StopProcessing()
bool StopProcessing() noexcept

Stops the active processor. Safe to call from inside a callback to break out of a blocking StartProcessing(). Returns bool.

IsOk()
bool IsOk(const FrameResult& result) noexcept

Returns bool — true when the window succeeded (result.error is empty).

AudioProcessorBuilder

SetSourceType()
SetSourceType(AudioSourceType)

Sets the source: kFile or kRecorder. Returns AudioProcessorBuilder& (chainable).

SetFilePath()
SetFilePath(const std::string&)

Sets the audio file path (required for kFile). Returns AudioProcessorBuilder&.

SetCallback()
SetCallback(FrameResultListener*)

Sets the result callback (required). Returns AudioProcessorBuilder&.

Build()
Build(std::string* error = nullptr)

Builds the processor. Returns std::unique_ptr<AudioProcessor> — nullptr on failure (message in *error).

AudioProcessor

StartProcessing()
bool StartProcessing(std::string* error = nullptr)

File mode runs to completion (blocks); stream mode arms the stream. Returns bool (message in *error on failure).

PushAudioChunk()
bool PushAudioChunk(const uint8_t* data, int size, int num_channels, SampleFormat format, int input_sample_rate)

Feeds one chunk of captured audio (stream mode); the SDK buffers it internally and runs inference only once a full model window has accumulated — so a single push may deliver zero, one, or several results.

ParameterTypeDescription
dataconst uint8_t*Interleaved PCM for one hop.
sizeintByte count of data.
num_channelsintChannel count (≥ 1).
formatSampleFormatEncoding of data.
input_sample_rateintCapture rate in Hz; must be ≥ model_sample_rate().

Returns bool — true if the chunk was accepted.

AudioProcessor::StopProcessing()
bool StopProcessing()

Stops this processor. Returns bool.

model_sample_rate()
int model_sample_rate()

Returns int — the model’s native sample rate in Hz.

controls()
Controls& controls()

Returns Controls& — the runtime feature controls (see Runtime controls below).

Runtime controls

Runtime feature controls on the processor (obtained from processor->controls()). AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".

Setter / GetterTypeDescription
SetSensitivity(level) / GetSensitivity()std::stringGlobal detection sensitivity.
SetTagSensitivity(tag, level) / GetTagSensitivity(tag)std::stringPer-tag sensitivity.
SetResultSummaryEnabled(enable) / IsResultSummaryEnabled()boolMerge same-tag windows into intervals.
SetResultSummaryMargin(seconds) / GetResultSummaryMargin()intInterval merge margin (seconds).
EnableAudioActivityDetection(enable) / IsAudioActivityDetectionEnabled()boolSkip inference on inactive audio.
SetAadSensitivity(level) / GetAadSensitivity()std::stringActivity-detection sensitivity.
SetAadHistoryCount(seconds) / GetAadHistoryCount()intActivity window length (seconds).
EnableAutomaticGainControl(enable) / IsAutomaticGainControlEnabled()boolNormalize input loudness.

FrameResult

Delivered to FrameResultListener::OnResult. IsOk(result) reports whether it succeeded.

FieldTypeDescription
tagsstd::vector<Tag>Detected tags; each Tag has name (std::string), probability (float).
start_time / end_timefloatWindow bounds, in seconds.
prediction_time_msdoubleProcessing time, in milliseconds.
timestampint64_tWindow timestamp.
summariesstd::vector<std::string>Interval lines (present when Result Summary is enabled).
errorstd::stringError message (empty on success).

FrameResultListener

Subclass and override void OnResult(const FrameResult&). It fires once per inference window on the SDK’s inference thread.

SampleFormat

kFloat32 | kInt16 | kInt32 | kFloat64 — the encoding passed to PushAudioChunk. Feed 24-bit audio as kInt32, left-justified (sample << 8).

7. Removed in v1.6.1

The pre-1.6 source-object API was replaced by the builder/processor API above. Migrate as follows:

RemovedUse instead
sense::Init(key, config_path) (returned int)sense::Init(key, config_contents, &err) (returns bool)
sense::AudioSourceFile + .Load() / .Predict()AudioProcessorBuilder (kFile) + StartProcessing()
sense::AudioSourceStream + .Predict(buffer, rate)AudioProcessorBuilder (kRecorder) + PushAudioChunk()
sense::Result (returned by Predict)sense::FrameResult (delivered to FrameResultListener::OnResult)

8. Troubleshooting

If you hit an error, reset the SDK’s local state and run again:

sudo rm -rf ~/.sense-sdk

1. Requirements

The Cochl.Sense Edge SDK for Android supports Android API 26 (8.0 ‘Oreo’) or later, and API 34+ with NDK r26. Android Studio is required. Beyond the shared git unzip curl tools, no extra apt packages are needed. Unlike C++/Python, Android needs no license daemon — the native license runs in-process and is set up automatically by the SDK.

2. Prepare the sample

(1) Clone the tutorial

git clone https://github.com/cochlearai/sense-sdk-android-tutorials.git

(2) Download the SDK

Download the AAR package built for your NDK (use ndk-r22b if you build with the older NDK r22):

curl -LO https://github.com/cochlearai/sense-sdk-android-tutorials/releases/download/v1.6.1/sense-sdk-1.6.1-android-ndk-r26b.zip

(3) Unzip the SDK into the tutorial directory

Unzip into each sample’s app module so the AAR lands in app/libs/:

# sense-file
unzip sense-sdk-1.6.1-android-ndk-r26b.zip \
  -d sense-sdk-android-tutorials/sense-file/app
# sense-stream
unzip sense-sdk-1.6.1-android-ndk-r26b.zip \
  -d sense-sdk-android-tutorials/sense-stream/app

(4) Modify Gradle and Manifest files

  • Add the AAR to the dependencies block in app/build.gradle:

    dependencies {
        implementation files('libs/sense-sdk-v1.6.1-ndk-r26b.aar')  // Cochl.Sense Edge SDK
    }
    
  • Add the required permissions to app/src/main/AndroidManifest.xml:

    <uses-permission android:name="android.permission.INTERNET" />
    <uses-permission android:name="android.permission.RECORD_AUDIO" />  <!-- for sense-stream -->
    

3. How to use the Cochl.Sense Edge SDK for Android

(1) Initialization

Initialize the SDK with your project key and the path to config.json. Initialization may take some time if a model must be downloaded.

import ai.cochl.sensesdk.CochlException;
import ai.cochl.sensesdk.Sense;

private final String projectKey = "Your project key";  // Replace with your own key
private final String configPath = "config/config.json";

try {
    sense = Sense.getInstance();
    File configFile = new File(this.getExternalFilesDir(null), configPath);
    sense.init(projectKey, configFile.getAbsolutePath());
} catch (CochlException e) {
    // Init has failed.
}

(2) Configuration

This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.

{
  "device_name": "SenseDevice",
  "num_threads": 0,
  "log_level": "info",
  "app_data_path": "~",
  "default_hopsize": 1.0,
  "metrics": {
    "retention_period": 0,
    "free_disk_space": 100,
    "push_period": 30
  },
  "sensitivity_control": {
    "default_sensitivity": "NORMAL",
    "tag_sensitivity": {
    }
  },
  "result_summary": {
    "enable": false,
    "default_interval_margin": 0,
    "tag_interval_margin": {
    }
  },
  "audio_preprocessor": {
    "audio_activity_detection": {
      "enable": false,
      "history_count": 600,
      "sensitivity": "NORMAL"
    },
    "automatic_gain_control": {
      "enable": false
    }
  }
}

Top-level fields

FieldDefaultDescription
device_namerequiredName shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page.
num_threads0Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1).
log_level"info"Log verbosity, as a string: "error", "warn", "info", "debug".
app_data_path"~"Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it.
default_hopsize1.0Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0].
metricsobjectLocal metrics buffer + push behavior. See below.
sensitivity_controlobjectDetection sensitivity, globally and per tag. See below and Advanced Configurations.
result_summaryobjectMerge consecutive same-tag windows into single interval events. See below and Advanced Configurations.
audio_preprocessorobjectAudio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations.

metrics sub-fields

Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.

FieldDefaultUnitDescription
retention_period0daysHow long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days.
free_disk_space100MBWhen free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB.
push_period30secondsHow often metrics are pushed to the dashboard. Capped at 3,600 s.

sensitivity_control sub-fields

Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".

FieldDefaultDescription
default_sensitivity"NORMAL"Detection sensitivity applied to every tag.
tag_sensitivity{}Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity.

result_summary sub-fields

When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).

FieldDefaultUnitDescription
enablefalse—Turns Result Summary on.
default_interval_margin0secondsGap tolerance for merging same-tag windows into one interval.
tag_interval_margin{}secondsPer-tag margin overrides, e.g. { "Speech": 2 }.

audio_preprocessor sub-fields

Both preprocessors apply to stream mode only.

FieldDefaultDescription
audio_activity_detection.enablefalseSkips inference on windows with no audio activity.
audio_activity_detection.history_count600Detection window length, in seconds (a fixed span, independent of hop size).
audio_activity_detection.sensitivity"NORMAL"Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH".
automatic_gain_control.enablefalseNormalizes input loudness before inference.

(3) Audio input and predict

The SDK receives audio and returns detected sound tags as JSON. Pass either a file path (string) or an audio-sample array. predict() returns { "frames": [ ... ] } — one entry per inference window completed during the call. With an array, the SDK buffers what you push: frames is [] until a full window has accumulated, holds several when you push more than one window’s worth, and keeps the remainder for the next call.

Sample rate. The input must be at or above the model’s rate (typically 22,050 Hz). Lower rates are not supported — file mode rejects a below-rate file with a resampling error and stream mode rejects the pushed chunk; higher rates are automatically downsampled. Use Sense.getInstance().modelSampleRate() to capture at the model’s native rate.

Pass the file path as a string.

Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.

import java.io.File;
import org.json.JSONObject;

import ai.cochl.sensesdk.Sense;

File file = new File(this.getExternalFilesDir(null), "audio_file.wav");
JSONObject result = Sense.getInstance().predict(file.getAbsolutePath());

Pass an audio-sample array and its sample rate. Push one hop of newly captured audio per call — the SDK buffers and windows internally.

Note: grant the microphone (RECORD_AUDIO) permission at runtime before capturing, and make sure an input device is available. This tutorial uses android.media.AudioRecord; you can replace it with another source.

import android.media.AudioFormat;
import android.media.AudioRecord;
import android.media.MediaRecorder;
import org.json.JSONObject;

import ai.cochl.sensesdk.Sense;

final int SAMPLE_RATE = 22050;   // or Sense.getInstance().modelSampleRate()
// Supported formats: PCM_FLOAT (float[]), PCM_16BIT (short[]).
final int AUDIO_FORMAT = AudioFormat.ENCODING_PCM_FLOAT;
final int bufSize = AudioRecord.getMinBufferSize(
    SAMPLE_RATE, AudioFormat.CHANNEL_IN_MONO, AUDIO_FORMAT);

AudioRecord recorder = new AudioRecord(
    MediaRecorder.AudioSource.UNPROCESSED, SAMPLE_RATE,
    AudioFormat.CHANNEL_IN_MONO, AUDIO_FORMAT, bufSize);
recorder.startRecording();

// Read one hop of audio into 'buffer' (float[] or short[]), then:
JSONObject frameResult = Sense.getInstance().predict(buffer, SAMPLE_RATE);

(4) Result format

Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.

predict() returns the completed windows rather than invoking a callback you register — the SDK holds the listener internally and batches the results into one object. frames is [] when you pushed less than a full window and holds several when you pushed more, with the remainder buffered for your next call:

{
  "frames": [
    {
      "start_time": 3.0,
      "end_time": 5.0,
      "prediction_time_ms": 24.0,
      "timestamp": 1730000000,
      "tags": [{ "name": "Siren", "probability": 0.81 }],
      "summaries": [],
      "error": ""
    }
  ]
}

Each entry of frames carries these fields:

FieldTypeNotes
tagsarray of objectsDetected tags; empty for silence or sub-threshold audio — a valid, successful result.
start_timenumberSeconds from the start of the stream or file.
end_timenumberstart_time plus the window length.
prediction_time_msnumberInference latency for this window.
timestampnumber (integer)Unix time, in seconds, when the window was processed.
summariesarray of stringsInterval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval.
errorstringEmpty on success; non-empty only on failure, with tags then empty.

Check error rather than testing tags — an empty tags means silence, not an error.

(5) Terminate

The SDK allocates resources during initialization. Release them with terminate().

import ai.cochl.sensesdk.Sense;

Sense.getInstance().terminate();

4. Build and run

Open each sample (sense-file, sense-stream) in Android Studio, set your project key in MainActivity, connect a device or start an emulator, and run the app. For sense-stream, accept the microphone permission prompt on first launch.

5. Reference

The SDK is a singleton (ai.cochl.sensesdk.Sense) driven by a project key. After a successful init, call predict and the runtime controls below.

Sense

getInstance()
static Sense getInstance()

Returns Sense — the singleton instance.

init()
void init(String projectKey, String configFilePath)

Authenticates with the project key and applies the configuration. Throws CochlException on failure.

ParameterTypeDescription
projectKeyStringYour project key.
configFilePathStringPath to config.json.
predict()
JSONObject predict(...)

Runs inference and returns { "frames": [ ... ] } — one entry per window completed during this call, so it may be empty or hold several. Throws CochlException on failure. Three overloads:

OverloadInput
predict(String filePath)Audio file (WAV or MP3).
predict(short[] shortArr, int sampleRate)16-bit PCM samples + capture rate (Hz).
predict(float[] floatArr, int sampleRate)32-bit float samples + capture rate (Hz).

Returns JSONObject.

modelSampleRate()
int modelSampleRate()

Returns int — the model’s native sample rate in Hz.

stopProcessing()
boolean stopProcessing()

Stops the current stream processing. Returns boolean.

isInitialized()
boolean isInitialized()

Returns boolean — whether the SDK is initialized.

terminate()
void terminate()

Releases the resources allocated during initialization.

Runtime controls

Runtime feature controls on the singleton. AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".

Setter / GetterTypeDescription
setSensitivity(level) / getSensitivity()StringGlobal detection sensitivity.
setTagSensitivity(tag, level) / getTagSensitivity(tag)StringPer-tag sensitivity.
setResultSummaryEnabled(enable) / isResultSummaryEnabled()booleanMerge same-tag windows into intervals.
setResultSummaryMargin(seconds) / getResultSummaryMargin()intInterval merge margin (seconds).
enableAudioActivityDetection(enable) / isAudioActivityDetectionEnabled()booleanSkip inference on inactive audio.
setAadSensitivity(level) / getAadSensitivity()StringActivity-detection sensitivity.
setAadHistoryCount(seconds) / getAadHistoryCount()intActivity window length (seconds).
enableAutomaticGainControl(enable) / isAutomaticGainControlEnabled()booleanNormalize input loudness.

CochlException

Thrown by init and predict on failure. Extends java.lang.RuntimeException.

6. Removed in v1.6.1

The following methods were removed. Migrate as follows:

RemovedUse instead
predict(byte[] byteArr, int sampleRate)predict(short[], int) or predict(float[], int)
getWindowSize() / getHopSize()config.json default_hopsize; the SDK windows internally
getSelectedTags() / getParameters()— (query the project on the dashboard)
getSdkVersion()—
addInput(AudioRecord) / addInput(File)predict(...)
predict(Sense.OnPredictListener)predict(...) returning JSONObject
pause() / resume() / stopPredict()stopProcessing()