Basic Feature Extraction
This guide shows you how to perform basic feature extraction from electrophysiological data using a probe insertion ID (PID).
Prerequisites
ibleatools installed and configured
Access to electrophysiological data
A valid probe insertion ID (PID)
Overview
The basic feature extraction workflow involves: 1. Loading data using a PID 2. Computing features from the raw electrophysiological signals 3. Saving the results for further analysis
Step-by-Step Guide
Import Required Modules
from pathlib import Path from one.api import ONE from ephysatlas.feature_computation import compute_features_from_pid
Set Up Parameters
# Define your PID and parameters pid = "your-probe-insertion-id-here" output_dir = Path("/path/to/output/directory") # Optional: specify time range and snippet durations t_start = 300.0 # Start time in seconds duration_ap = 1.0 # AP snippet length in seconds duration_lf = 1.0 # LF snippet length in seconds
Run Feature Extraction
# Compute features (pass a ONE client; it loads the raw data) one = ONE() result = compute_features_from_pid( pid=pid, one=one, output_dir=output_dir, t_start=t_start, duration_ap=duration_ap, duration_lf=duration_lf, ) print(f"Features computed successfully!") print(f"Output saved to: {output_dir}")
Verify Results
# Check what was created output_path = output_dir / pid if output_path.exists(): print(f"Output directory created: {output_path}") # List generated files for item in output_path.rglob("*"): if item.is_file(): print(f" {item.relative_to(output_path)}")
Complete Example
Here’s the complete script from examples/feature_extraction_example.py:
"""Basic feature extraction from a probe insertion ID (PID).
Computes AP / LF / CSD features for a short snippet using the public
``compute_features_from_pid`` entry point, which loads the raw data through ONE,
destripes it, computes features, and writes them under ``output_dir``.
Run as a script or cell-by-cell (``# %%`` cells).
"""
# %% Imports and configuration
import logging
import tempfile
from pathlib import Path
from one.api import ONE
from ephysatlas.feature_calculators import CsdParams, FeatureParams
from ephysatlas.feature_computation import compute_features_from_pid
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
# A valid probe insertion ID and where to write the outputs.
pid = "4cb60c5c-d15b-4abd-8cfd-776bc5a81dbe"
output_dir = Path(tempfile.mkdtemp(prefix="ephysatlas_features_"))
logger.info("Writing features under %s", output_dir)
one = ONE()
# %% Compute features for a 1-second snippet
# duration_ap / duration_lf set the AP and LF snippet lengths independently
# (the older single ``duration`` argument is deprecated).
df = compute_features_from_pid(
pid=pid,
one=one,
t_start=300.0,
duration_ap=1.0,
duration_lf=1.0,
features_to_compute=["lf", "csd", "ap"],
output_dir=output_dir,
)
logger.info("Computed features: shape=%s, columns=%s", df.shape, sorted(df.columns))
# %% Optional: override per-feature parameters
# ``feature_params`` accepts the typed objects OR a nested dict; here we turn off
# CSD scaling. Only the options you set change; everything else keeps its default.
df_unscaled_csd = compute_features_from_pid(
pid=pid,
one=one,
t_start=300.0,
duration_ap=1.0,
duration_lf=1.0,
features_to_compute=["csd"],
output_dir=output_dir,
feature_params=FeatureParams(csd=CsdParams(scale=False)),
# equivalently: feature_params={"csd": {"scale": False}}
)
logger.info("CSD (scale=False) features: shape=%s", df_unscaled_csd.shape)
Customizing per-feature parameters
Per-feature options are passed via feature_params — either the typed objects
from ephysatlas.feature_calculators or an equivalent nested dict. Only the
options you set change; everything else keeps its default. For example, to
disable CSD scaling:
from ephysatlas.feature_calculators import FeatureParams, CsdParams
result = compute_features_from_pid(
pid=pid,
one=one,
t_start=t_start,
duration_ap=duration_ap,
duration_lf=duration_lf,
feature_params=FeatureParams(csd=CsdParams(scale=False)),
# equivalently: feature_params={"csd": {"scale": False}}
)
Using the OOP calculators directly
compute_features_from_pid and compute_features_from_file are thin
wrappers over the calculators in ephysatlas.feature_calculators. Use a
calculator directly when you also want the intermediate destriped snippet (for
inspection or plotting), or when working from local SpikeGLX files:
from ephysatlas.feature_calculators import (
IBLPIDFeatureCalculator,
FeatureComputationOptions,
SnippetWindow,
)
calc = IBLPIDFeatureCalculator(pid=pid, one=one)
window = SnippetWindow(t_start=300.0, duration_ap=1.0, duration_lf=1.0)
options = FeatureComputationOptions(
features_to_compute=["lf", "csd", "ap"], output_dir=output_dir
)
result = calc.compute_snippet(window, options)
# Intermediate destriped data, without recomputing features:
snippet = calc.get_destriped_snippet(window)
For local files use
ephysatlas.feature_calculators.SpikeGLXFileFeatureCalculator (or
ephysatlas.feature_computation.compute_features_from_file()). The full OOP
script is in examples/feature_extraction_oop.py.
Expected Output
When successful, you should see: * A new directory structure created under your output directory * Feature files in various formats (.pqt, .npy) * Log messages indicating successful completion * No error messages
Troubleshooting
Common Issues:
PID not found: Ensure the PID exists in your data repository
Permission errors: Check write permissions for the output directory
Memory issues: For large datasets, consider processing smaller time chunks
Missing dependencies: Ensure all required packages are installed
Getting Help:
Check the logs for detailed error messages
Verify your data access permissions
Consult the API Reference for detailed API documentation