Samples
Two samples are provided in the Cochl.Sense Edge SDK: sense-file and sense-stream.
sense-fileruns inference on an audio file (WAV or MP3).sense-streamruns inference on a live audio buffer (e.g. a microphone).
Make sure your environment is set up first — see Getting Started. Every platform below shares the same three preparation steps — clone the tutorial, download the SDK, unzip the SDK into the tutorial — and needs these command-line tools. On Linux hosts:
sudo apt update
sudo apt install -y git unzip curl
Then pick the platform you’re targeting. Each tab follows the same order: Requirements → Prepare → License manager → Usage → Build & run → Reference → Removed in v1.6.1 → Troubleshooting. Android has no License manager step (the license runs in-process) and no Troubleshooting section.
1. Requirements
The Cochl.Sense Edge SDK for Python supports Python 3.6–3.12. The wheel is a compiled, version-specific build, so you must download the file that matches both your Python version and CPU architecture (e.g. cp38 for Python 3.8, aarch64 for a 64-bit Raspberry Pi). If you need a version that is not listed, contact support@cochl.ai.
In addition to the shared git unzip curl tools above, install the Python build tools and the PortAudio library the sense-stream sample uses for microphone capture:
sudo apt install -y python3-venv python3-pip libportaudio2
Check your interpreter version first — you’ll download the matching SDK in step 2:
python3 --version # e.g. Python 3.8.10 -> download the cp38 wheel
2. Prepare the sample
(1) Clone the tutorial
git clone https://github.com/cochlearai/sense-sdk-python-tutorials.git
(2) Create and activate a virtual environment
Create the environment with the interpreter whose version matches the SDK you’re about to download:
cd sense-sdk-python-tutorials
python3 -m venv ./venv
source ./venv/bin/activate # prompt is now prefixed with (venv)
python -m pip install --upgrade pip
- NOTE: For older Python versions (e.g. 3.6, 3.7) the venv module ships in a version-specific package (
python3.6-venv,python3.7-venv). Install the one matching your interpreter.
(3) Download the version-matched SDK
Download the wheel package built for your Python version and architecture. Replace cp38 and linux_aarch64 with the values matching your environment — note Python 3.6 and 3.7 use the ABI tags cp36-cp36m / cp37-cp37m, not cp36-cp36 / cp37-cp37. Resources lists the exact link for every version:
# (venv)
curl -LO https://github.com/cochlearai/sense-sdk-python-tutorials/releases/download/v1.6.1/sense-1.6.1-cp38-cp38-linux_aarch64.zip
(4) Unzip the SDK into the tutorial directory
unzip sense-1.6.1-cp38-cp38-linux_aarch64.zip -d .
After unzipping, your sample directory should look like this:
sense-sdk-python-tutorials
|-- README.md
|-- config.json # configuration file used to initialize the SDK
|-- audio-files # audio samples
|-- sense-file # sense-file sample (sense_file.py)
|-- sense-stream # sense-stream sample (sense_stream.py)
`-- sense # Cochl.Sense Edge SDK
|-- sense-1.6.1-cp38-cp38-linux_aarch64.whl
`-- sl # license daemon (scripts + binary)
(5) Install the SDK and dependencies
Install the wheel you unzipped. The Sense wheel has no runtime dependencies; the sense-stream sample additionally needs sounddevice (from its requirements.txt):
# (venv)
pip install sense/sense-1.6.1-cp38-cp38-linux_aarch64.whl
pip install -r sense-stream/requirements.txt # sense-stream only (sounddevice)
3. Install the license manager
The license manager is required from SDK v1.4.0. From v1.6.1 it is the native sense-lm backend, which replaced Thales. On Linux it runs as a small system daemon that verifies the license and needs an internet connection. Install it once:
sudo ./sense/sl/scripts/install_daemon.sh
The installer creates a system group sense-lm and adds your user to it. Apply the new group membership without logging out:
newgrp sense-lm # or log out and back in
To remove the daemon later:
sudo ./sense/sl/scripts/uninstall_daemon.sh
4. How to use the Cochl.Sense Edge SDK for Python
(1) Initialization
Initialize the SDK with your project key and the contents of config.json. sense.session(...) is a context manager that calls init() on entry and terminate() on exit. Initialization may take some time if a model must be downloaded (the SDK downloads a model when none is present or a newer one is available).
import sense
project_key = "Your project key" # Replace with your own key
with open("config.json", encoding="utf-8") as f:
config = f.read()
with sense.session(project_key, config):
# ... create a processor and run inference here ...
pass
(2) Configuration
This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.
{
"device_name": "SenseDevice",
"num_threads": 0,
"log_level": "info",
"app_data_path": "~",
"default_hopsize": 1.0,
"metrics": {
"retention_period": 0,
"free_disk_space": 100,
"push_period": 30
},
"sensitivity_control": {
"default_sensitivity": "NORMAL",
"tag_sensitivity": {
}
},
"result_summary": {
"enable": false,
"default_interval_margin": 0,
"tag_interval_margin": {
}
},
"audio_preprocessor": {
"audio_activity_detection": {
"enable": false,
"history_count": 600,
"sensitivity": "NORMAL"
},
"automatic_gain_control": {
"enable": false
}
}
}
Top-level fields
| Field | Default | Description |
|---|---|---|
device_name | required | Name shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page. |
num_threads | 0 | Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1). |
log_level | "info" | Log verbosity, as a string: "error", "warn", "info", "debug". |
app_data_path | "~" | Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it. |
default_hopsize | 1.0 | Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0]. |
metrics | object | Local metrics buffer + push behavior. See below. |
sensitivity_control | object | Detection sensitivity, globally and per tag. See below and Advanced Configurations. |
result_summary | object | Merge consecutive same-tag windows into single interval events. See below and Advanced Configurations. |
audio_preprocessor | object | Audio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations. |
metrics sub-fields
Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.
| Field | Default | Unit | Description |
|---|---|---|---|
retention_period | 0 | days | How long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days. |
free_disk_space | 100 | MB | When free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB. |
push_period | 30 | seconds | How often metrics are pushed to the dashboard. Capped at 3,600 s. |
sensitivity_control sub-fields
Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".
| Field | Default | Description |
|---|---|---|
default_sensitivity | "NORMAL" | Detection sensitivity applied to every tag. |
tag_sensitivity | {} | Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity. |
result_summary sub-fields
When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).
| Field | Default | Unit | Description |
|---|---|---|---|
enable | false | — | Turns Result Summary on. |
default_interval_margin | 0 | seconds | Gap tolerance for merging same-tag windows into one interval. |
tag_interval_margin | {} | seconds | Per-tag margin overrides, e.g. { "Speech": 2 }. |
audio_preprocessor sub-fields
Both preprocessors apply to stream mode only.
| Field | Default | Description |
|---|---|---|
audio_activity_detection.enable | false | Skips inference on windows with no audio activity. |
audio_activity_detection.history_count | 600 | Detection window length, in seconds (a fixed span, independent of hop size). |
audio_activity_detection.sensitivity | "NORMAL" | Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH". |
automatic_gain_control.enable | false | Normalizes input loudness before inference. |
(3) Audio input and predict
The SDK receives audio and returns detected sound tags. In file mode you point it at a file; in stream mode you push captured audio buffers. Results are delivered to a sense.ResultListener callback (on_result), one call per inference window.
processor.model_sample_rate() to capture at the model’s native rate.Create a file processor and start it. start() blocks until the whole file is processed.
Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.
import sense
class Printer(sense.ResultListener):
def on_result(self, frame: sense.Result) -> None:
if not frame.is_ok():
print("error:", frame.error)
return
for tag in frame.tags:
print(frame.start_time, tag.name, tag.probability)
listener = Printer()
processor = sense.create_file_processor("audio_file.wav", listener)
processor.start() # blocks until end of file
Create a stream processor, then push captured audio chunks. Use the processor as a context manager (with processor:) to arm/stop the stream.
Note: connect a microphone (or other input device) before running the stream sample. This tutorial uses sounddevice (PortAudio) for capture; you can replace it with any source.
import sounddevice as sd
import sense
listener = Printer() # sense.ResultListener
processor = sense.create_stream_processor(listener)
rate = processor.model_sample_rate() # capture at the model rate
mic = sd.RawInputStream(samplerate=rate, channels=1, dtype="int16")
with mic, processor: # processor.start() arms the stream
while True:
data, _ = mic.read(rate) # one hop of audio
processor.push(bytes(data), 1, sense.SampleFormat.INT16, rate)
(4) Result format
Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.
Your ResultListener.on_result is called once per window with a sense.Result:
class Tag:
name: str # Detected tag, e.g. "Siren"
probability: float # Confidence for that tag, 0.0-1.0
class Result:
tags: List[Tag] # Empty when nothing passed threshold
start_time: float # Window start, seconds
end_time: float # Window end, seconds
prediction_time_ms: float # Processing time for this window, ms
timestamp: int # Unix time (seconds) the window was processed
summaries: List[str] # Result Summary lines; empty unless enabled
error: str # Empty on success, message on failure
def is_ok(self) -> bool: ... # True when error is empty
| Field | Type | Notes |
|---|---|---|
tags | List[Tag] | Detected tags; empty for silence or sub-threshold audio — a valid, successful result. |
start_time | float | Seconds from the start of the stream or file. |
end_time | float | start_time plus the window length. |
prediction_time_ms | float | Inference latency for this window. |
timestamp | int | Unix time, in seconds, when the window was processed. |
summaries | List[str] | Interval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval. |
error | str | Empty on success; non-empty only on failure, with tags then empty. |
Check result.is_ok() rather than testing tags — an empty tags means silence, not an error.
(5) Terminate
The SDK allocates resources during initialization. Release them with terminate(). When you use sense.session(...), this happens automatically on exit; call it explicitly only if you used sense.init(...) directly.
import sense
sense.terminate()
5. How to run
Run from the tutorial root so the samples find config.json in the working directory. Set your project key in PROJECT_KEY at the top of sense-file/sense_file.py and sense-stream/sense_stream.py first.
# (venv)
python sense-file/sense_file.py <PATH_TO_AUDIO_FILE>
python sense-file/sense_file.py audio-files/babycry.wav
Note: connect a microphone before running.
# (venv)
python sense-stream/sense_stream.py # Ctrl-C to stop
6. Reference
Import from the sense package.
Module functions
sense.init()
sense.init(project_key, config_json)
Authenticates with the project key and applies the configuration.
| Parameter | Type | Description |
|---|---|---|
project_key | str | Your project key. |
config_json | str | The contents of config.json (not a path). |
Returns bool. Raises RuntimeError on failure — it does not return False, so if not sense.init(...) never fires. Wrap calls in try / except RuntimeError, or use with sense.session(...), which does it for you.
sense.terminate()
sense.terminate()
Releases all resources allocated during initialization. Returns bool — True on success.
sense.is_initialized()
sense.is_initialized()
Returns bool — whether the SDK is initialized.
sense.session()
sense.session(project_key, config_json)
Context manager that calls init() on entry and terminate() on exit: with sense.session(key, config): .... Parameters are the same as init. Returns a context manager.
sense.create_file_processor()
sense.create_file_processor(file_path, listener)
Creates a file-mode processor.
| Parameter | Type | Description |
|---|---|---|
file_path | str | Path to the audio file (WAV or MP3). |
listener | ResultListener | Receives one on_result call per window. |
Returns Processor. Raises RuntimeError on failure.
sense.create_stream_processor()
sense.create_stream_processor(listener)
Creates a stream-mode processor fed by push.
| Parameter | Type | Description |
|---|---|---|
listener | ResultListener | Receives one on_result call per window. |
Returns Processor. Raises RuntimeError on failure.
Processor
Processor.start()
Processor.start()
File mode runs to completion (blocks); stream mode arms the stream. Returns bool. Raises RuntimeError on failure.
Processor.push()
Processor.push(pcm, num_channels, sample_format, input_sample_rate)
Feeds one chunk of captured audio (stream mode); the SDK buffers it internally and runs inference only once a full model window has accumulated — so a single push may deliver zero, one, or several results.
| Parameter | Type | Description |
|---|---|---|
pcm | bytes | Interleaved PCM for one hop. |
num_channels | int | Channel count (≥ 1). |
sample_format | int | A SampleFormat value. |
input_sample_rate | int | Capture rate in Hz; must be ≥ model_sample_rate(). |
Returns bool — True if the chunk was accepted.
Processor.stop()
Processor.stop()
Stops processing. Returns bool.
Processor.model_sample_rate()
Processor.model_sample_rate()
Returns int — the model’s native sample rate in Hz.
Runtime controls
Runtime feature controls on the processor. AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".
| Setter / Getter | Type | Description |
|---|---|---|
set_sensitivity(level) / get_sensitivity() | str | Global detection sensitivity. |
set_tag_sensitivity(tag, level) / get_tag_sensitivity(tag) | str | Per-tag sensitivity. |
set_result_summary_enabled(enable) / is_result_summary_enabled() | bool | Merge same-tag windows into intervals. |
set_result_summary_margin(seconds) / get_result_summary_margin() | int | Interval merge margin (seconds). |
enable_audio_activity_detection(enable) / is_audio_activity_detection_enabled() | bool | Skip inference on inactive audio. |
set_aad_sensitivity(level) / get_aad_sensitivity() | str | Activity-detection sensitivity. |
set_aad_history_count(seconds) / get_aad_history_count() | int | Activity window length (seconds). |
enable_automatic_gain_control(enable) / is_automatic_gain_control_enabled() | bool | Normalize input loudness. |
ResultListener
Subclass and override on_result(self, frame: Result) -> None. It fires once per inference window on the SDK’s inference thread, so keep the callback cheap.
Result
Passed to on_result. is_ok() -> bool reports whether the window succeeded.
| Field | Type | Description |
|---|---|---|
tags | List[Tag] | Detected tags; each Tag has name: str, probability: float. |
start_time / end_time | float | Window bounds, in seconds. |
prediction_time_ms | float | Processing time, in milliseconds. |
timestamp | int | Window timestamp. |
summaries | List[str] | Interval lines (present when Result Summary is enabled). |
error | str | Error message (empty when is_ok()). |
SampleFormat
SampleFormat.FLOAT32 | INT16 | INT32 | FLOAT64 — the encoding passed to push. FLOAT64 is for file/pre-recorded input only.
7. Removed in v1.6.1
The pre-1.6 procedural API was replaced by the processor/listener API above. Migrate as follows:
| Removed | Use instead |
|---|---|
SenseInit(project_key, config_path) | sense.init(project_key, config_contents) or sense.session(...) |
SenseTerminate() | sense.terminate() |
AudioSourceFile().Load() + .Predict() | sense.create_file_processor(path, listener) + processor.start() |
AudioSourceStream().Predict(buffer, rate) | sense.create_stream_processor(listener) + processor.push(...) |
8. Troubleshooting
If you hit an error, reset the SDK’s local state and run again:
sudo rm -rf ~/.sense-sdk
1. Requirements
In addition to the shared git unzip curl tools above, the C++ sample needs a C++14 compiler and CMake (≥ 3.13, for the cmake -B option used below):
sudo apt install -y build-essential cmake
2. Prepare the sample
(1) Clone the tutorial
git clone https://github.com/cochlearai/sense-sdk-cpp-tutorials.git
(2) Download the SDK
Download the Linux build for your architecture (replace x86_64 with aarch64 on 64-bit Arm):
cd sense-sdk-cpp-tutorials
curl -LO https://github.com/cochlearai/sense-sdk-cpp-tutorials/releases/download/v1.6.1/sense-sdk-1.6.1-linux-x86_64-cpp.zip
(3) Unzip the SDK into the tutorial directory
unzip sense-sdk-1.6.1-linux-x86_64-cpp.zip -d .
After unzipping, your sample directory should look like this:
sense-sdk-cpp-tutorials
|-- CMakeLists.txt
|-- README.md
|-- config.json # configuration file used to initialize the SDK
|-- audio-files # audio samples
|-- sense-file # sense-file sample (sense_file.cc)
|-- sense-stream # sense-stream sample (sense_stream.cc)
`-- sense # Cochl.Sense Edge SDK
|-- include
|-- lib
`-- sl # license daemon (scripts + binary)
3. Install the license manager
The license manager is required from SDK v1.4.0. From v1.6.1 it is the native sense-lm backend, which replaced Thales. On Linux it runs as a small system daemon that verifies the license and needs an internet connection. Install it once:
sudo ./sense/sl/scripts/install_daemon.sh
The installer creates a system group sense-lm and adds your user to it. Apply the new group membership without logging out:
newgrp sense-lm # or log out and back in
To remove the daemon later:
sudo ./sense/sl/scripts/uninstall_daemon.sh
4. How to use the Cochl.Sense Edge SDK for C++
(1) Initialization
Initialize the SDK with your project key and the contents of config.json (read the file yourself and pass the string). sense::Init returns false on failure and writes a message to the optional error pointer. Initialization may take some time if a model must be downloaded.
#include <string>
#include "sense/sense.hpp"
std::string project_key = "Your project key"; // Replace with your own key
std::string config = /* contents of config.json */;
std::string err;
if (!sense::Init(project_key, config, &err)) {
// Init has failed; see err.
return -1;
}
(2) Configuration
This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.
{
"device_name": "SenseDevice",
"num_threads": 0,
"log_level": "info",
"app_data_path": "~",
"default_hopsize": 1.0,
"metrics": {
"retention_period": 0,
"free_disk_space": 100,
"push_period": 30
},
"sensitivity_control": {
"default_sensitivity": "NORMAL",
"tag_sensitivity": {
}
},
"result_summary": {
"enable": false,
"default_interval_margin": 0,
"tag_interval_margin": {
}
},
"audio_preprocessor": {
"audio_activity_detection": {
"enable": false,
"history_count": 600,
"sensitivity": "NORMAL"
},
"automatic_gain_control": {
"enable": false
}
}
}
Top-level fields
| Field | Default | Description |
|---|---|---|
device_name | required | Name shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page. |
num_threads | 0 | Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1). |
log_level | "info" | Log verbosity, as a string: "error", "warn", "info", "debug". |
app_data_path | "~" | Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it. |
default_hopsize | 1.0 | Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0]. |
metrics | object | Local metrics buffer + push behavior. See below. |
sensitivity_control | object | Detection sensitivity, globally and per tag. See below and Advanced Configurations. |
result_summary | object | Merge consecutive same-tag windows into single interval events. See below and Advanced Configurations. |
audio_preprocessor | object | Audio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations. |
metrics sub-fields
Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.
| Field | Default | Unit | Description |
|---|---|---|---|
retention_period | 0 | days | How long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days. |
free_disk_space | 100 | MB | When free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB. |
push_period | 30 | seconds | How often metrics are pushed to the dashboard. Capped at 3,600 s. |
sensitivity_control sub-fields
Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".
| Field | Default | Description |
|---|---|---|
default_sensitivity | "NORMAL" | Detection sensitivity applied to every tag. |
tag_sensitivity | {} | Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity. |
result_summary sub-fields
When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).
| Field | Default | Unit | Description |
|---|---|---|---|
enable | false | — | Turns Result Summary on. |
default_interval_margin | 0 | seconds | Gap tolerance for merging same-tag windows into one interval. |
tag_interval_margin | {} | seconds | Per-tag margin overrides, e.g. { "Speech": 2 }. |
audio_preprocessor sub-fields
Both preprocessors apply to stream mode only.
| Field | Default | Description |
|---|---|---|
audio_activity_detection.enable | false | Skips inference on windows with no audio activity. |
audio_activity_detection.history_count | 600 | Detection window length, in seconds (a fixed span, independent of hop size). |
audio_activity_detection.sensitivity | "NORMAL" | Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH". |
automatic_gain_control.enable | false | Normalizes input loudness before inference. |
(3) Audio input and predict
Build an AudioProcessor with a source type and a FrameResultListener. Each inference window is delivered to your listener’s OnResult callback.
processor->model_sample_rate() to capture at the model’s native rate.#include "sense/sense.hpp"
class Printer : public sense::FrameResultListener {
public:
void OnResult(const sense::FrameResult& result) override {
if (!sense::IsOk(result)) { /* handle result.error */ return; }
for (const auto& tag : result.tags) { /* tag.name, tag.probability */ }
}
};
Build a kFile processor and start it. StartProcessing() blocks until the whole file is processed.
Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.
#include <memory>
#include "sense/sense.hpp"
Printer listener;
std::string build_err;
std::unique_ptr<sense::AudioProcessor> processor =
sense::AudioProcessorBuilder()
.SetSourceType(sense::AudioSourceType::kFile)
.SetFilePath("audio_file.wav")
.SetCallback(&listener)
.Build(&build_err);
if (!processor) { /* handle build_err */ }
std::string start_err;
processor->StartProcessing(&start_err); // blocks until end of file
Build a kRecorder processor, arm it with StartProcessing(), then push captured audio via PushAudioChunk().
Note: connect a microphone (or other input device) before running the stream sample. The tutorial captures audio itself (vendored miniaudio); the SDK has no audio-device dependency.
#include <memory>
#include "sense/sense.hpp"
Printer listener;
std::unique_ptr<sense::AudioProcessor> processor =
sense::AudioProcessorBuilder()
.SetSourceType(sense::AudioSourceType::kRecorder)
.SetCallback(&listener)
.Build();
processor->StartProcessing(); // arms the stream
const int rate = processor->model_sample_rate();
// For each captured chunk (interleaved PCM), push it to the SDK:
processor->PushAudioChunk(
reinterpret_cast<const uint8_t*>(buffer), size_in_bytes,
/*num_channels=*/1, sense::SampleFormat::kFloat32, rate);
(4) Result format
Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.
Your FrameResultListener::OnResult is called once per window with a sense::FrameResult:
struct FrameResult {
struct Tag {
std::string name; // Detected tag, e.g. "Siren"
float probability; // Confidence for that tag, 0.0-1.0
};
std::vector<Tag> tags; // Empty when nothing passed threshold
float start_time; // Window start, seconds
float end_time; // Window end, seconds
double prediction_time_ms; // Processing time for this window, ms
int64_t timestamp; // Unix time (seconds) the window was processed
std::vector<std::string> summaries; // Result Summary lines; empty unless enabled
std::string error; // Empty on success, message on failure
};
| Field | Type | Notes |
|---|---|---|
tags | std::vector<Tag> | Detected tags; empty for silence or sub-threshold audio — a valid, successful result. |
start_time | float | Seconds from the start of the stream or file. |
end_time | float | start_time plus the window length. |
prediction_time_ms | double | Inference latency for this window. |
timestamp | int64_t | Unix time, in seconds, when the window was processed. |
summaries | std::vector<std::string> | Interval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval. |
error | std::string | Empty on success; non-empty only on failure, with tags then empty. |
Check IsOk(result) rather than testing tags — an empty tags means silence, not an error.
(5) Terminate
The SDK allocates resources during initialization. Release them with sense::Terminate().
#include "sense/sense.hpp"
sense::Terminate();
5. Build and run
The tutorial ships a CMakeLists.txt. Set your project key in kProjectKey at the top of sense-file/sense_file.cc and sense-stream/sense_stream.cc, then build:
cmake -B build
cmake --build build
This produces sense-file-app and sense-stream-app next to config.json and audio-files/, so you can run them in place.
./sense-file-app <PATH_TO_AUDIO_FILE>
./sense-file-app audio-files/babycry.wav
Note: connect a microphone before running.
./sense-stream-app # Ctrl-C to stop
6. Reference
Declared in sense/sense.hpp (namespace sense).
Free functions
Init()
bool Init(const std::string& project_key, std::string config_json, std::string* error = nullptr)
Authenticates with the project key and applies the configuration.
| Parameter | Type | Description |
|---|---|---|
project_key | const std::string& | Your project key. |
config_json | std::string | The contents of config.json (not a path). |
error | std::string* | Optional; receives a message on failure. |
Returns bool — true on success.
Terminate()
bool Terminate() noexcept
Releases all resources allocated during initialization. Returns bool.
IsInitialized()
bool IsInitialized() noexcept
Returns bool — whether the SDK is initialized.
sense::StopProcessing()
bool StopProcessing() noexcept
Stops the active processor. Safe to call from inside a callback to break out of a blocking StartProcessing(). Returns bool.
IsOk()
bool IsOk(const FrameResult& result) noexcept
Returns bool — true when the window succeeded (result.error is empty).
AudioProcessorBuilder
SetSourceType()
SetSourceType(AudioSourceType)
Sets the source: kFile or kRecorder. Returns AudioProcessorBuilder& (chainable).
SetFilePath()
SetFilePath(const std::string&)
Sets the audio file path (required for kFile). Returns AudioProcessorBuilder&.
SetCallback()
SetCallback(FrameResultListener*)
Sets the result callback (required). Returns AudioProcessorBuilder&.
Build()
Build(std::string* error = nullptr)
Builds the processor. Returns std::unique_ptr<AudioProcessor> — nullptr on failure (message in *error).
AudioProcessor
StartProcessing()
bool StartProcessing(std::string* error = nullptr)
File mode runs to completion (blocks); stream mode arms the stream. Returns bool (message in *error on failure).
PushAudioChunk()
bool PushAudioChunk(const uint8_t* data, int size, int num_channels, SampleFormat format, int input_sample_rate)
Feeds one chunk of captured audio (stream mode); the SDK buffers it internally and runs inference only once a full model window has accumulated — so a single push may deliver zero, one, or several results.
| Parameter | Type | Description |
|---|---|---|
data | const uint8_t* | Interleaved PCM for one hop. |
size | int | Byte count of data. |
num_channels | int | Channel count (≥ 1). |
format | SampleFormat | Encoding of data. |
input_sample_rate | int | Capture rate in Hz; must be ≥ model_sample_rate(). |
Returns bool — true if the chunk was accepted.
AudioProcessor::StopProcessing()
bool StopProcessing()
Stops this processor. Returns bool.
model_sample_rate()
int model_sample_rate()
Returns int — the model’s native sample rate in Hz.
controls()
Controls& controls()
Returns Controls& — the runtime feature controls (see Runtime controls below).
Runtime controls
Runtime feature controls on the processor (obtained from processor->controls()). AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".
| Setter / Getter | Type | Description |
|---|---|---|
SetSensitivity(level) / GetSensitivity() | std::string | Global detection sensitivity. |
SetTagSensitivity(tag, level) / GetTagSensitivity(tag) | std::string | Per-tag sensitivity. |
SetResultSummaryEnabled(enable) / IsResultSummaryEnabled() | bool | Merge same-tag windows into intervals. |
SetResultSummaryMargin(seconds) / GetResultSummaryMargin() | int | Interval merge margin (seconds). |
EnableAudioActivityDetection(enable) / IsAudioActivityDetectionEnabled() | bool | Skip inference on inactive audio. |
SetAadSensitivity(level) / GetAadSensitivity() | std::string | Activity-detection sensitivity. |
SetAadHistoryCount(seconds) / GetAadHistoryCount() | int | Activity window length (seconds). |
EnableAutomaticGainControl(enable) / IsAutomaticGainControlEnabled() | bool | Normalize input loudness. |
FrameResult
Delivered to FrameResultListener::OnResult. IsOk(result) reports whether it succeeded.
| Field | Type | Description |
|---|---|---|
tags | std::vector<Tag> | Detected tags; each Tag has name (std::string), probability (float). |
start_time / end_time | float | Window bounds, in seconds. |
prediction_time_ms | double | Processing time, in milliseconds. |
timestamp | int64_t | Window timestamp. |
summaries | std::vector<std::string> | Interval lines (present when Result Summary is enabled). |
error | std::string | Error message (empty on success). |
FrameResultListener
Subclass and override void OnResult(const FrameResult&). It fires once per inference window on the SDK’s inference thread.
SampleFormat
kFloat32 | kInt16 | kInt32 | kFloat64 — the encoding passed to PushAudioChunk. Feed 24-bit audio as kInt32, left-justified (sample << 8).
7. Removed in v1.6.1
The pre-1.6 source-object API was replaced by the builder/processor API above. Migrate as follows:
| Removed | Use instead |
|---|---|
sense::Init(key, config_path) (returned int) | sense::Init(key, config_contents, &err) (returns bool) |
sense::AudioSourceFile + .Load() / .Predict() | AudioProcessorBuilder (kFile) + StartProcessing() |
sense::AudioSourceStream + .Predict(buffer, rate) | AudioProcessorBuilder (kRecorder) + PushAudioChunk() |
sense::Result (returned by Predict) | sense::FrameResult (delivered to FrameResultListener::OnResult) |
8. Troubleshooting
If you hit an error, reset the SDK’s local state and run again:
sudo rm -rf ~/.sense-sdk
1. Requirements
The Cochl.Sense Edge SDK for Android supports Android API 26 (8.0 ‘Oreo’) or later, and API 34+ with NDK r26. Android Studio is required. Beyond the shared git unzip curl tools, no extra apt packages are needed. Unlike C++/Python, Android needs no license daemon — the native license runs in-process and is set up automatically by the SDK.
2. Prepare the sample
(1) Clone the tutorial
git clone https://github.com/cochlearai/sense-sdk-android-tutorials.git
(2) Download the SDK
Download the AAR package built for your NDK (use ndk-r22b if you build with the older NDK r22):
curl -LO https://github.com/cochlearai/sense-sdk-android-tutorials/releases/download/v1.6.1/sense-sdk-1.6.1-android-ndk-r26b.zip
(3) Unzip the SDK into the tutorial directory
Unzip into each sample’s app module so the AAR lands in app/libs/:
# sense-file
unzip sense-sdk-1.6.1-android-ndk-r26b.zip \
-d sense-sdk-android-tutorials/sense-file/app
# sense-stream
unzip sense-sdk-1.6.1-android-ndk-r26b.zip \
-d sense-sdk-android-tutorials/sense-stream/app
(4) Modify Gradle and Manifest files
Add the AAR to the
dependenciesblock inapp/build.gradle:dependencies { implementation files('libs/sense-sdk-v1.6.1-ndk-r26b.aar') // Cochl.Sense Edge SDK }Add the required permissions to
app/src/main/AndroidManifest.xml:<uses-permission android:name="android.permission.INTERNET" /> <uses-permission android:name="android.permission.RECORD_AUDIO" /> <!-- for sense-stream -->
3. How to use the Cochl.Sense Edge SDK for Android
(1) Initialization
Initialize the SDK with your project key and the path to config.json. Initialization may take some time if a model must be downloaded.
import ai.cochl.sensesdk.CochlException;
import ai.cochl.sensesdk.Sense;
private final String projectKey = "Your project key"; // Replace with your own key
private final String configPath = "config/config.json";
try {
sense = Sense.getInstance();
File configFile = new File(this.getExternalFilesDir(null), configPath);
sense.init(projectKey, configFile.getAbsolutePath());
} catch (CochlException e) {
// Init has failed.
}
(2) Configuration
This file stores the Cochl.Sense Edge SDK’s configuration settings. Passing its contents (C++/Python) or its path (Android) to the initialization method lets the SDK read and apply the settings. All fields are optional except device_name; omitted fields fall back to the defaults below.
{
"device_name": "SenseDevice",
"num_threads": 0,
"log_level": "info",
"app_data_path": "~",
"default_hopsize": 1.0,
"metrics": {
"retention_period": 0,
"free_disk_space": 100,
"push_period": 30
},
"sensitivity_control": {
"default_sensitivity": "NORMAL",
"tag_sensitivity": {
}
},
"result_summary": {
"enable": false,
"default_interval_margin": 0,
"tag_interval_margin": {
}
},
"audio_preprocessor": {
"audio_activity_detection": {
"enable": false,
"history_count": 600,
"sensitivity": "NORMAL"
},
"automatic_gain_control": {
"enable": false
}
}
}
Top-level fields
| Field | Default | Description |
|---|---|---|
device_name | required | Name shown for this device on the dashboard. One device_name per device. Change it later via the Edge SDK tab of your project page. |
num_threads | 0 | Threads used for inference. A positive value is the exact count; 0 uses all cores; a negative value counts down from all cores — -1 = all cores, -2 = all cores minus one, -3 = all cores minus two, and so on (result clamped to at least 1). |
log_level | "info" | Log verbosity, as a string: "error", "warn", "info", "debug". |
app_data_path | "~" | Where the SDK stores resources generated during execution (downloaded models, metrics buffer, license state). "~" expands to the user’s home directory. On Android this is set automatically to the app’s private storage — you don’t need to configure it. |
default_hopsize | 1.0 | Inference hop in seconds — how often a result is produced. Clamped to [0.1, 1.0]. |
metrics | object | Local metrics buffer + push behavior. See below. |
sensitivity_control | object | Detection sensitivity, globally and per tag. See below and Advanced Configurations. |
result_summary | object | Merge consecutive same-tag windows into single interval events. See below and Advanced Configurations. |
audio_preprocessor | object | Audio Activity Detection + Automatic Gain Control (stream mode only). See below and Advanced Configurations. |
metrics sub-fields
Each inference window produces a metric, buffered locally and pushed to the dashboard on a schedule.
| Field | Default | Unit | Description |
|---|---|---|---|
retention_period | 0 | days | How long the SDK keeps unsent metrics locally. 0 disables local saving. Capped at 31 days. |
free_disk_space | 100 | MB | When free disk drops below this threshold, the SDK stops storing metrics locally (effectively retention_period: 0). Capped at 1,000,000 MB. |
push_period | 30 | seconds | How often metrics are pushed to the dashboard. Capped at 3,600 s. |
sensitivity_control sub-fields
Sensitivity levels are strings: "VERY_LOW", "LOW", "NORMAL", "HIGH", "VERY_HIGH".
| Field | Default | Description |
|---|---|---|
default_sensitivity | "NORMAL" | Detection sensitivity applied to every tag. |
tag_sensitivity | {} | Per-tag overrides, e.g. { "Glass_break": "VERY_HIGH", "Rustle": "LOW" }. Tags not listed use default_sensitivity. |
result_summary sub-fields
When enabled, consecutive windows of the same tag are merged into one interval event (summaries in the result).
| Field | Default | Unit | Description |
|---|---|---|---|
enable | false | — | Turns Result Summary on. |
default_interval_margin | 0 | seconds | Gap tolerance for merging same-tag windows into one interval. |
tag_interval_margin | {} | seconds | Per-tag margin overrides, e.g. { "Speech": 2 }. |
audio_preprocessor sub-fields
Both preprocessors apply to stream mode only.
| Field | Default | Description |
|---|---|---|
audio_activity_detection.enable | false | Skips inference on windows with no audio activity. |
audio_activity_detection.history_count | 600 | Detection window length, in seconds (a fixed span, independent of hop size). |
audio_activity_detection.sensitivity | "NORMAL" | Activity-detection sensitivity: "VERY_LOW" … "VERY_HIGH". |
automatic_gain_control.enable | false | Normalizes input loudness before inference. |
(3) Audio input and predict
The SDK receives audio and returns detected sound tags as JSON. Pass either a file path (string) or an audio-sample array. predict() returns { "frames": [ ... ] } — one entry per inference window completed during the call. With an array, the SDK buffers what you push: frames is [] until a full window has accumulated, holds several when you push more than one window’s worth, and keeps the remainder for the next call.
Sense.getInstance().modelSampleRate() to capture at the model’s native rate.Pass the file path as a string.
Note: a file shorter than the model’s window is still processed — the SDK zero-pads the final window; only an empty file is rejected.
import java.io.File;
import org.json.JSONObject;
import ai.cochl.sensesdk.Sense;
File file = new File(this.getExternalFilesDir(null), "audio_file.wav");
JSONObject result = Sense.getInstance().predict(file.getAbsolutePath());
Pass an audio-sample array and its sample rate. Push one hop of newly captured audio per call — the SDK buffers and windows internally.
Note: grant the microphone (RECORD_AUDIO) permission at runtime before capturing, and make sure an input device is available. This tutorial uses android.media.AudioRecord; you can replace it with another source.
import android.media.AudioFormat;
import android.media.AudioRecord;
import android.media.MediaRecorder;
import org.json.JSONObject;
import ai.cochl.sensesdk.Sense;
final int SAMPLE_RATE = 22050; // or Sense.getInstance().modelSampleRate()
// Supported formats: PCM_FLOAT (float[]), PCM_16BIT (short[]).
final int AUDIO_FORMAT = AudioFormat.ENCODING_PCM_FLOAT;
final int bufSize = AudioRecord.getMinBufferSize(
SAMPLE_RATE, AudioFormat.CHANNEL_IN_MONO, AUDIO_FORMAT);
AudioRecord recorder = new AudioRecord(
MediaRecorder.AudioSource.UNPROCESSED, SAMPLE_RATE,
AudioFormat.CHANNEL_IN_MONO, AUDIO_FORMAT, bufSize);
recorder.startRecording();
// Read one hop of audio into 'buffer' (float[] or short[]), then:
JSONObject frameResult = Sense.getInstance().predict(buffer, SAMPLE_RATE);
(4) Result format
Each inference window produces one result. In stream mode, pushing audio only appends it to an internal FIFO; it does not produce a result for each call. When the buffer contains a full model window, the SDK runs inference, drops the oldest hop, and repeats as long as a full window remains. Pushing less than one window produces no results, while pushing several windows’ worth produces several results. Any leftover audio stays buffered for the next push. In file mode, the SDK produces one result per window across the entire file.
predict() returns the completed windows rather than invoking a callback you register — the SDK holds the listener internally and batches the results into one object. frames is [] when you pushed less than a full window and holds several when you pushed more, with the remainder buffered for your next call:
{
"frames": [
{
"start_time": 3.0,
"end_time": 5.0,
"prediction_time_ms": 24.0,
"timestamp": 1730000000,
"tags": [{ "name": "Siren", "probability": 0.81 }],
"summaries": [],
"error": ""
}
]
}
Each entry of frames carries these fields:
| Field | Type | Notes |
|---|---|---|
tags | array of objects | Detected tags; empty for silence or sub-threshold audio — a valid, successful result. |
start_time | number | Seconds from the start of the stream or file. |
end_time | number | start_time plus the window length. |
prediction_time_ms | number | Inference latency for this window. |
timestamp | number (integer) | Unix time, in seconds, when the window was processed. |
summaries | array of strings | Interval lines; only with Result Summary on. Stream mode emits "Listening..." for a window with no interval. |
error | string | Empty on success; non-empty only on failure, with tags then empty. |
Check error rather than testing tags — an empty tags means silence, not an error.
(5) Terminate
The SDK allocates resources during initialization. Release them with terminate().
import ai.cochl.sensesdk.Sense;
Sense.getInstance().terminate();
4. Build and run
Open each sample (sense-file, sense-stream) in Android Studio, set your project key in MainActivity, connect a device or start an emulator, and run the app. For sense-stream, accept the microphone permission prompt on first launch.
5. Reference
The SDK is a singleton (ai.cochl.sensesdk.Sense) driven by a project key. After a successful init, call predict and the runtime controls below.
Sense
getInstance()
static Sense getInstance()
Returns Sense — the singleton instance.
init()
void init(String projectKey, String configFilePath)
Authenticates with the project key and applies the configuration. Throws CochlException on failure.
| Parameter | Type | Description |
|---|---|---|
projectKey | String | Your project key. |
configFilePath | String | Path to config.json. |
predict()
JSONObject predict(...)
Runs inference and returns { "frames": [ ... ] } — one entry per window completed during this call, so it may be empty or hold several. Throws CochlException on failure. Three overloads:
| Overload | Input |
|---|---|
predict(String filePath) | Audio file (WAV or MP3). |
predict(short[] shortArr, int sampleRate) | 16-bit PCM samples + capture rate (Hz). |
predict(float[] floatArr, int sampleRate) | 32-bit float samples + capture rate (Hz). |
Returns JSONObject.
modelSampleRate()
int modelSampleRate()
Returns int — the model’s native sample rate in Hz.
stopProcessing()
boolean stopProcessing()
Stops the current stream processing. Returns boolean.
isInitialized()
boolean isInitialized()
Returns boolean — whether the SDK is initialized.
terminate()
void terminate()
Releases the resources allocated during initialization.
Runtime controls
Runtime feature controls on the singleton. AAD/AGC apply to stream mode only. Sensitivity levels are "VERY_LOW" | "LOW" | "NORMAL" | "HIGH" | "VERY_HIGH".
| Setter / Getter | Type | Description |
|---|---|---|
setSensitivity(level) / getSensitivity() | String | Global detection sensitivity. |
setTagSensitivity(tag, level) / getTagSensitivity(tag) | String | Per-tag sensitivity. |
setResultSummaryEnabled(enable) / isResultSummaryEnabled() | boolean | Merge same-tag windows into intervals. |
setResultSummaryMargin(seconds) / getResultSummaryMargin() | int | Interval merge margin (seconds). |
enableAudioActivityDetection(enable) / isAudioActivityDetectionEnabled() | boolean | Skip inference on inactive audio. |
setAadSensitivity(level) / getAadSensitivity() | String | Activity-detection sensitivity. |
setAadHistoryCount(seconds) / getAadHistoryCount() | int | Activity window length (seconds). |
enableAutomaticGainControl(enable) / isAutomaticGainControlEnabled() | boolean | Normalize input loudness. |
CochlException
Thrown by init and predict on failure. Extends java.lang.RuntimeException.
6. Removed in v1.6.1
The following methods were removed. Migrate as follows:
| Removed | Use instead |
|---|---|
predict(byte[] byteArr, int sampleRate) | predict(short[], int) or predict(float[], int) |
getWindowSize() / getHopSize() | config.json default_hopsize; the SDK windows internally |
getSelectedTags() / getParameters() | — (query the project on the dashboard) |
getSdkVersion() | — |
addInput(AudioRecord) / addInput(File) | predict(...) |
predict(Sense.OnPredictListener) | predict(...) returning JSONObject |
pause() / resume() / stopPredict() | stopProcessing() |