Features Module

The features module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings.

Module Overview

This module includes: * Feature extraction from AP (action potential) and LF (local field potential) bands * Current source density (CSD) computation * Spike detection and waveform analysis * Feature denoising using total variation filters * Data validation using Pandera schemas * Transformer classes for scikit-learn compatibility

Core Functions

Feature Computation

Electrophysiological feature extraction and processing module.

This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.

The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility

Classes

DartParameters

Configuration parameters for Dartsort backend

ChannelDataFrameSchema

Pandera schema for channel data validation

ModelLfFeatures

Schema for local field potential features

ModelCsdFeatures

Schema for current source density features

ModelApFeatures

Schema for action potential features

ModelSpikeFeatures

Schema for spike waveform features

ModelSpikeShapeFeatures

Sparse, neuroscientist-facing remapping of ModelSpikeFeatures

ModelChannelLayout

Schema for channel layout information

ModelHistologyPlanned

Schema for planned histology coordinates

ModelHistologyResolved

Schema for resolved histology coordinates

ModelRawFeatures

Combined schema for all raw features

EphysTransformer

Transformer for applying feature transformations

EphysDenoiser

Transformer for denoising electrophysiological features

Functions

_setup_scratch_directory

Set up scratch directory with fallback logic

voltage_features_set

Get list of feature column names by provenance

lf

Compute LF features from numpy array

csd

Compute CSD features from numpy array

ap

Compute AP features from numpy array

dart_subtraction_numpy

Perform spike detection using Dartsort

spikes

Spike detection and feature extraction with multiple backend support

remap_waveform_shape_features

Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set

remap_waveform_shape_features_volume

Same remapping, for a brainwide encoding volume instead of a dataframe

xcor_acor_ratio

Compute cross-correlation over auto-correlation ratio

denoise_shank

Denoise AP features using total variation filter

denoise_dataframe

Apply total variation filter denoising to features

Constants

__features_version__str

Version of the feature extractor code

BANDSdict

Frequency bands for spectral analysis

FEATURES_LISTlist

List of available feature types

Examples

>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000)  # 64 channels, 30k samples
>>> fs = 30000  # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))

Notes

This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.

See Also

ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities

ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]

Set up scratch directory with fallback logic.

Parameters:

scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.

Returns:

Path to the created scratch directory.

Return type:

Path

Note

This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.

ephysatlas.features.get_feature_cmin(feature_name)[source]

Get the minimum value for a given feature.

Parameters:

feature_name (str) – Name of the feature.

Returns:

Minimum value for the feature.

Return type:

float

Note

This function is currently a placeholder and needs implementation.

class ephysatlas.features.DartParameters(**data)[source]

Bases: BaseModel

Configuration parameters for Dartsort backend.

This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.

Variables:
  • localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.

  • chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).

  • trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.

  • scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.

  • n_jobs (int) – DARTsort worker count used to build its ComputationConfig. 0 runs in the main process; positive values select that many workers. Defaults to 0.

  • detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as voltage_threshold. Defaults to 4.0.

  • spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.

  • positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.

  • residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as subtraction_threshold. Defaults to 3.162 (sqrt(10)).

localization_radius: Annotated[float, Gt(gt=0)]
chunk_length_samples: Annotated[int, Gt(gt=0)]
trough_offset: Annotated[int, Gt(gt=0)]
detection_threshold: Annotated[float, Gt(gt=0)]
spatial_dedup_radius_um: Annotated[float, Gt(gt=0)]
positive_temporal_dedup_radius_samples: Annotated[int, Gt(gt=0)]
residnorm_decrease_threshold: Annotated[float, Gt(gt=0)]
scratch_dir: Path | str | None
n_jobs: Annotated[int, Ge(ge=0)]
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]

Bases: DataFrameModel

Pandera schema for channel data validation.

This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.

Note

x/y/z are in metres, while axial_um/lateral_um are in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.

Variables:
  • pid (Series[str]) – Probe insertion ID.

  • channel (Series[int]) – Channel index.

  • x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).

  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

  • acronym (Series[str]) – Brain region acronym.

  • atlas_id (Series[int]) – Atlas region identifier.

pid: str = 'pid'
channel: int = 'channel'
x: float = 'x'
y: float = 'y'
z: float = 'z'
axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
acronym: str = 'acronym'
atlas_id: int = 'atlas_id'
distance_to_tip_um: float | None = 'distance_to_tip_um'
class Config

Bases: BaseConfig

name: str | None = 'ChannelDataFrameSchema'

name of schema

class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]

Bases: DataFrameModel

Base class for channel-based feature schemas.

This is an abstract base class that provides the foundation for all channel-based feature validation schemas.

Note

The channel field is expected to be an index in derived classes.

class Config

Bases: BaseConfig

name: str | None = 'BaseChannelFeatures'

name of schema

class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for local field potential features.

This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf (Series[float]) – Root mean square of LFP signal in dB.

  • psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).

  • psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).

  • psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).

  • psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).

  • psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).

  • psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).

rms_lf: float = 'rms_lf'
psd_lfp: float = 'psd_lfp'
psd_delta: float = 'psd_delta'
psd_theta: float = 'psd_theta'
psd_alpha: float = 'psd_alpha'
psd_beta: float = 'psd_beta'
psd_gamma: float = 'psd_gamma'
psd_residual_lfp: float | None = 'psd_residual_lfp'
psd_residual_delta: float | None = 'psd_residual_delta'
psd_residual_theta: float | None = 'psd_residual_theta'
psd_residual_alpha: float | None = 'psd_residual_alpha'
psd_residual_beta: float | None = 'psd_residual_beta'
psd_residual_gamma: float | None = 'psd_residual_gamma'
aperiodic_offset: float | None = 'aperiodic_offset'
aperiodic_exponent: float | None = 'aperiodic_exponent'
decay_fit_error: float | None = 'decay_fit_error'
decay_fit_r_squared: float | None = 'decay_fit_r_squared'
decay_n_peaks: float | None = 'decay_n_peaks'
rms_lf_no_car: float | None = 'rms_lf_no_car'
class Config

Bases: Config

name: str | None = 'ModelLfFeatures'

name of schema

class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for current source density features.

This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.

  • psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).

  • psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).

  • psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).

  • psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).

  • psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).

  • psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).

rms_lf_csd: float = 'rms_lf_csd'
psd_delta_csd: float = 'psd_delta_csd'
psd_theta_csd: float = 'psd_theta_csd'
psd_alpha_csd: float = 'psd_alpha_csd'
psd_beta_csd: float = 'psd_beta_csd'
psd_gamma_csd: float = 'psd_gamma_csd'
psd_lfp_csd: float = 'psd_lfp_csd'
rms_lf_csd_diff1: float | None = 'rms_lf_csd_diff1'
psd_delta_csd_diff1: float | None = 'psd_delta_csd_diff1'
psd_theta_csd_diff1: float | None = 'psd_theta_csd_diff1'
psd_alpha_csd_diff1: float | None = 'psd_alpha_csd_diff1'
psd_beta_csd_diff1: float | None = 'psd_beta_csd_diff1'
psd_gamma_csd_diff1: float | None = 'psd_gamma_csd_diff1'
psd_lfp_csd_diff1: float | None = 'psd_lfp_csd_diff1'
class Config

Bases: Config

name: str | None = 'ModelCsdFeatures'

name of schema

class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for action potential features.

This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.

Variables:
  • rms_ap (Series[float]) – Root mean square of AP signal in dB.

  • cor_ratio (Series[float]) – Cross-correlation over auto-correlation ratio.

  • channel_labels (Series[int]) – Quality labels for channels.

rms_ap: float = 'rms_ap'
cor_ratio: float = 'cor_ratio'
channel_labels: int = 'channel_labels'
class Config

Bases: Config

name: str | None = 'ModelApFeatures'

name of schema

class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Spike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).

Variables:
  • alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).

  • alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).

  • depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).

  • polarity (Series[float]) – Spike polarity, dimensionless.

  • recovery_slope (Series[float]) – Slope during recovery phase (V/s).

  • repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).

  • slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).

  • slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).

  • spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).

  • spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).

  • spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).

  • tip_val (Series[float]) – Tip amplitude value (V).

Note

tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.

alpha_mean: float = 'alpha_mean'
alpha_std: float = 'alpha_std'
depolarisation_slope: float = 'depolarisation_slope'
polarity: float = 'polarity'
recovery_slope: float = 'recovery_slope'
repolarisation_slope: float = 'repolarisation_slope'
slowness_s_per_m: float | None = 'slowness_s_per_m'
slowness_s_per_m_std: float | None = 'slowness_s_per_m_std'
spatial_spread_um: float | None = 'spatial_spread_um'
spatial_spread_um_std: float | None = 'spatial_spread_um_std'
spike_count: float = 'spike_count'
tip_val: float = 'tip_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeSharedFeatures'

name of schema

class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for spike waveform features.

This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.

Variables:
  • peak_time_secs (Series[float]) – Time to peak in seconds.

  • peak_val (Series[float]) – Peak amplitude value (V).

  • recovery_time_secs (Series[float]) – Recovery time in seconds.

  • tip_time_secs (Series[float]) – Time to tip in seconds.

  • trough_time_secs (Series[float]) – Time to trough in seconds.

  • trough_val (Series[float]) – Trough amplitude value (V).

Note

peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.

peak_time_secs: float = 'peak_time_secs'
peak_val: float = 'peak_val'
recovery_time_secs: float = 'recovery_time_secs'
tip_time_secs: float = 'tip_time_secs'
trough_time_secs: float = 'trough_time_secs'
trough_val: float = 'trough_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeFeatures'

name of schema

class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for the sparse, neuroscientist-facing spike-shape feature set.

Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.

The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.

Variables:
  • spike_width_secs (Series[float]) – Trough-to-peak spike duration (s).

  • predepolarisation_width_secs (Series[float]) – Tip-to-peak duration (s).

  • spike_amplitude (Series[float]) – Trough-to-peak amplitude (V).

  • peak_to_trough_ratio_log (Series[float]) – Log amplitude ratio (dimensionless).

spike_width_secs: float = 'spike_width_secs'
predepolarisation_width_secs: float = 'predepolarisation_width_secs'
spike_amplitude: float = 'spike_amplitude'
peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
class Config

Bases: Config

name: str | None = 'ModelSpikeShapeFeatures'

name of schema

ephysatlas.features._remap_waveform_shape_arrays(get)[source]

Core of the waveform-shape remapping, agnostic to the container.

Parameters

getcallable

get(name) returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).

Returns

dict

Maps each ModelSpikeShapeFeatures column name to its array.

ephysatlas.features.remap_waveform_shape_features(df)[source]

Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.

The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only recovery_time_secs (exact constant offset of trough_time_secs) and peak_val/trough_val/peak_time_secs/ trough_time_secs/tip_time_secs (reparametrised into spike_amplitude/peak_to_trough_ratio_log/spike_width_secs/ predepolarisation_width_secs without losing relative information) are actually dropped. tip_val and the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information. slowness_s_per_m/ spatial_spread_um (#123) and their _std counterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.

Return type:

DataFrame

Parameters

dfpandas.DataFrame

Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.

Returns

pandas.DataFrame

Columns of ModelSpikeShapeFeatures, same index as df.

ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]

Same transform as remap_waveform_shape_features, for a download_encoding_volume array.

Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.

Parameters

ephys_atlas_volnumpy.ndarray

Shape (nx, ny, nz, n_features).

feature_namesnumpy.ndarray or sequence of str

Length n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.

Returns

remapped_volnumpy.ndarray

Shape (nx, ny, nz, 16), dtype float32.

remapped_feature_namesnumpy.ndarray

Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).

Raises

KeyError

If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).

class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for channel layout information.

This schema defines the structure and validation rules for channel layout features including spatial positioning.

Variables:
  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
class Config

Bases: Config

name: str | None = 'ModelChannelLayout'

name of schema

class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for planned histology coordinates.

This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.

Note

Same coordinate space as ModelHistologyResolved (IBL, metres from Bregma); Populated by feature_computation.add_target_coordinates, which queries Alyx trajectories with provenance__lte,50 and takes the best provenance available per PROVENANCE_PREFERENCE: Micro-manipulator (30), else Planned (10), else Histology track (50); an insertion with none of the three raises ValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.

Variables:
  • x_target (Series[float]) – Target X-coordinate in metres from Bregma (IBL coordinates space).

  • y_target (Series[float]) – Target Y-coordinate in metres from Bregma (IBL coordinates space).

  • z_target (Series[float]) – Target Z-coordinate in metres from Bregma (IBL coordinates space).

x_target: float = 'x_target'
y_target: float = 'y_target'
z_target: float = 'z_target'
class Config

Bases: Config

name: str | None = 'ModelHistologyPlanned'

name of schema

class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for resolved histology coordinates.

This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.

Note

Same coordinate space as ModelHistologyPlanned. These come from SpikeSortingLoader.load_channels() – the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the *_target columns use, hence “at the most recent histology step”.

Variables:
  • x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).

  • atlas_id (Series[int]) – Atlas region identifier.

  • acronym (Series[str]) – Brain region acronym.

x: float = 'x'
y: float = 'y'
z: float = 'z'
atlas_id: int = 'atlas_id'
acronym: str = 'acronym'
class Config

Bases: Config

name: str | None = 'ModelHistologyResolved'

name of schema

class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]

Bases: ModelSpikeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for all raw features.

This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.

Note

This class inherits from multiple feature schemas to provide a unified interface for all feature types.

class Config

Bases: Config

name: str | None = 'ModelRawFeatures'

name of schema

class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]

Bases: ModelSpikeShapeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for denoised features, after the waveform-shape remap.

Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.

Note

Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.

class Config

Bases: Config

name: str | None = 'ModelDenoisedFeatures'

name of schema

class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]

Bases: DataFrameModel

Schema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.

pid: str = 'pid'
eid: str = 'eid'
probe_name: str | None = 'probe_name'
probe_serial: str | None = 'probe_serial'
neuropixel_version: str | None = 'neuropixel_version'
probe_model: str | None = 'probe_model'
referencing_scheme: str | None = 'referencing_scheme'
lab: str | None = 'lab'
pname: str | None = 'pname'
spike_sorting: str | None = 'spike_sorting'
spike_sorting_version: str | None = 'spike_sorting_version'
histology: str | None = 'histology'
record_length: float | None = 'record_length'
channel_count: float | None = 'channel_count'
bwm: bool = 'bwm'
class Config

Bases: BaseConfig

name: str | None = 'ModelProbeDetails'

name of schema

ephysatlas.features.feature_group_of_column_map()[source]

Map each denoisable feature column name to its provenance group.

Return type:

dict

Returns

dict

Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.

ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]

Get list of feature column names by provenance.

This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.

Parameters:

features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.

Returns:

Sorted list of feature column names excluding the ‘channel’ column.

Return type:

list

Note

The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.

ephysatlas.features._get_power_in_band(fscale, period, band)[source]

Calculate power in a specific frequency band.

Parameters:
  • fscale (np.ndarray) – Frequency scale array.

  • period (np.ndarray) – Periodogram values.

  • band (list) – Frequency band [low, high] in Hz.

Returns:

Power in the specified band in dB relative to v/sqrt(Hz).

Return type:

np.ndarray

Note

This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.

ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]

Extract power spectral density decay features from electrophysiological data.

This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.

Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.

The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.

Parameters

datanp.ndarray

Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.

fsfloat

Sampling frequency of the data in Hz.

fscalenp.ndarray

Frequency scale array corresponding to the period data.

periodnp.ndarray

2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.

bandsdict

Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.

npersegint, optional

Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.

PSD_rangelist of float, optional

Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.

Returns

pd.DataFrame

One row per channel. Aperiodic-component columns: aperiodic_offset (log10 power at 1 Hz), aperiodic_exponent (1/f slope in log-log space), decay_fit_error (RMS fit error), decay_fit_r_squared (goodness of fit) and decay_n_peaks (number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal): psd_residual_delta, psd_residual_theta, psd_residual_alpha, psd_residual_beta, psd_residual_gamma and psd_residual_lfp.

Notes

The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.

Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.

The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1

The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))

Examples

>>> import numpy as np
>>> # Generate sample data and frequency arrays
>>> data = np.random.randn(10, 10000)
>>> fs = 1000.0  # 1 kHz sampling rate
>>> fscale = np.linspace(0, 90, 100)
>>> period = np.random.rand(10, 100)  # Mock PSD data
>>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]}
>>> features = get_psd_decay_features(data, fs, fscale, period, bands)
>>> print(features.columns)
Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error',
       'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta',
       'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')

References

Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.

ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]

Compute LF features from a numpy array.

Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,

’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].

Return type:

pd.DataFrame

Note

The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.

ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]

Compute CSD features from a numpy array.

Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

  • decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10. 1 skips scipy.signal.decimate entirely (passed through unchanged) rather than calling it with a no-op factor: scipy.signal.decimate(..., q=1, ...) designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, which scipy.signal.firwin rejects.

  • scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.

  • denoise (bool, optional) – Apply ibldsp.cadzow.cadzow_denoiser to the decimated data before the CSD finite difference. Defaults to True (today’s behavior). Set False when data has already been Cadzow-denoised upstream (e.g. combined with decimate=1 for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,

’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].

Return type:

pd.DataFrame

Note

The function applies Cadzow denoising and current source density computation before computing the spectral features.

ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]

Compute AP features from a numpy array.

Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.

Parameters:
  • data (np.ndarray) – AP band data array with shape (channels, samples).

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • channel_labels (np.ndarray) – Array of channel quality labels.

Returns:

DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].

Return type:

pd.DataFrame

Raises:

AssertionError – If geometry or channel_labels are not provided.

Note

This function computes RMS values and cross-correlation ratios for action potential band data.

ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]

Perform spike detection using Dartsort.

This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.

Parameters:
  • data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • **params – Additional parameters for Dartsort configuration.

Returns:

A tuple containing:

  • df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.

  • d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.

Return type:

tuple

Note

This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.

DARTsort detects spikes on data normalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning: ptp and the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_val and the slopes computed by ibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.

ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]

Dartsort backend for spike detection.

This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Dartsort parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (DartParameters): Dartsort parameters object.

Return type:

tuple

ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]

SpikeInterface backend for spike detection.

This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – SpikeInterface parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (dict): SpikeInterface parameters object.

Return type:

tuple

Raises:

Note

This function is currently a placeholder and needs full implementation for SpikeInterface backend support.

ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]

Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.

Parameters:
  • channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in d_waveforms["channel_index"]. Padded with the sentinel n_channels_real for unused neighbour slots.

  • spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g. df_spikes["channel"].

  • x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.

  • y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.

Returns:

(n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused

neighbour slots.

Return type:

np.ndarray

ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]

Spike detection and feature extraction with multiple backend support.

This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.

  • return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.

  • backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Backend-specific parameters.

Returns:

If return_waveforms is False, returns DataFrame with

aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).

Return type:

pd.DataFrame or tuple

Raises:

ValueError – If an unknown backend is specified.

Note

The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.

ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]

Compute cross-correlation over auto-correlation ratio.

This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.

Parameters:
  • v (np.ndarray) – Voltage array for AP band with shape (nc, ns) where nc is number of channels and ns is number of samples.

  • geometry (dict) – Geometry dictionary with ‘x’ and ‘y’ arrays for electrode positions.

  • n_neighbor (int, optional) – Number of neighboring channels to consider. Defaults to 3.

Returns:

Array of size (nc,) containing the correlation ratios.

Return type:

np.ndarray

Note

The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.

ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]

Denoise AP features using total variation filter.

Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.

Parameters:
  • feature (np.ndarray) – AP feature to denoise with shape (nc,).

  • xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).

  • labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.

  • fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.

Returns:

Denoised AP features with shape (nc,).

Return type:

np.ndarray

Note

This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.

class ephysatlas.features._EphysTransformerInterface[source]

Bases: ABC, OneToOneFeatureMixin, TransformerMixin, BaseEstimator

Abstract base class for electrophysiological feature transformers.

This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.

Note

This is an abstract base class that should not be instantiated directly.

validate_X(X)[source]
Return type:

None

fit_transform(X=None, y=None)[source]

Fit the transformer and transform the data.

Parameters:
  • X (pd.DataFrame, optional) – Input data to fit and transform.

  • y – Ignored, present for compatibility with scikit-learn interface.

Returns:

Transformed data.

Return type:

pd.DataFrame

class ephysatlas.features.EphysTransformer[source]

Bases: _EphysTransformerInterface

fit(X=None, y=None)[source]
transform(X, y=None)[source]
class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Bases: _EphysTransformerInterface

__init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

TV-denoise electrophysiological features, with one weight factor per feature group.

Parameters

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.

channel_labelsnp.ndarray, optional

Channel quality annotation array with shape (nc,).

transform(X, y=None)[source]
fit(X=None, y=None)[source]
ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Applies total variation filter denoising to the features of a single probe insertion dataframe.

This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.

Parameters

df_pidpandas.DataFrame

DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.

Returns

pandas.DataFrame

A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.

Data Classes

Electrophysiological feature extraction and processing module.

This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.

The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility

Classes

DartParameters

Configuration parameters for Dartsort backend

ChannelDataFrameSchema

Pandera schema for channel data validation

ModelLfFeatures

Schema for local field potential features

ModelCsdFeatures

Schema for current source density features

ModelApFeatures

Schema for action potential features

ModelSpikeFeatures

Schema for spike waveform features

ModelSpikeShapeFeatures

Sparse, neuroscientist-facing remapping of ModelSpikeFeatures

ModelChannelLayout

Schema for channel layout information

ModelHistologyPlanned

Schema for planned histology coordinates

ModelHistologyResolved

Schema for resolved histology coordinates

ModelRawFeatures

Combined schema for all raw features

EphysTransformer

Transformer for applying feature transformations

EphysDenoiser

Transformer for denoising electrophysiological features

Functions

_setup_scratch_directory

Set up scratch directory with fallback logic

voltage_features_set

Get list of feature column names by provenance

lf

Compute LF features from numpy array

csd

Compute CSD features from numpy array

ap

Compute AP features from numpy array

dart_subtraction_numpy

Perform spike detection using Dartsort

spikes

Spike detection and feature extraction with multiple backend support

remap_waveform_shape_features

Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set

remap_waveform_shape_features_volume

Same remapping, for a brainwide encoding volume instead of a dataframe

xcor_acor_ratio

Compute cross-correlation over auto-correlation ratio

denoise_shank

Denoise AP features using total variation filter

denoise_dataframe

Apply total variation filter denoising to features

Constants

__features_version__str

Version of the feature extractor code

BANDSdict

Frequency bands for spectral analysis

FEATURES_LISTlist

List of available feature types

Examples

>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000)  # 64 channels, 30k samples
>>> fs = 30000  # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))

Notes

This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.

See Also

ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities

ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]

Set up scratch directory with fallback logic.

Parameters:

scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.

Returns:

Path to the created scratch directory.

Return type:

Path

Note

This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.

ephysatlas.features.get_feature_cmin(feature_name)[source]

Get the minimum value for a given feature.

Parameters:

feature_name (str) – Name of the feature.

Returns:

Minimum value for the feature.

Return type:

float

Note

This function is currently a placeholder and needs implementation.

class ephysatlas.features.DartParameters(**data)[source]

Bases: BaseModel

Configuration parameters for Dartsort backend.

This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.

Variables:
  • localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.

  • chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).

  • trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.

  • scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.

  • n_jobs (int) – DARTsort worker count used to build its ComputationConfig. 0 runs in the main process; positive values select that many workers. Defaults to 0.

  • detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as voltage_threshold. Defaults to 4.0.

  • spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.

  • positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.

  • residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as subtraction_threshold. Defaults to 3.162 (sqrt(10)).

localization_radius: Annotated[float, Gt(gt=0)]
chunk_length_samples: Annotated[int, Gt(gt=0)]
trough_offset: Annotated[int, Gt(gt=0)]
detection_threshold: Annotated[float, Gt(gt=0)]
spatial_dedup_radius_um: Annotated[float, Gt(gt=0)]
positive_temporal_dedup_radius_samples: Annotated[int, Gt(gt=0)]
residnorm_decrease_threshold: Annotated[float, Gt(gt=0)]
scratch_dir: Path | str | None
n_jobs: Annotated[int, Ge(ge=0)]
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]

Bases: DataFrameModel

Pandera schema for channel data validation.

This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.

Note

x/y/z are in metres, while axial_um/lateral_um are in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.

Variables:
  • pid (Series[str]) – Probe insertion ID.

  • channel (Series[int]) – Channel index.

  • x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).

  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

  • acronym (Series[str]) – Brain region acronym.

  • atlas_id (Series[int]) – Atlas region identifier.

pid: str = 'pid'
channel: int = 'channel'
x: float = 'x'
y: float = 'y'
z: float = 'z'
axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
acronym: str = 'acronym'
atlas_id: int = 'atlas_id'
distance_to_tip_um: float | None = 'distance_to_tip_um'
class Config

Bases: BaseConfig

name: str | None = 'ChannelDataFrameSchema'

name of schema

class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]

Bases: DataFrameModel

Base class for channel-based feature schemas.

This is an abstract base class that provides the foundation for all channel-based feature validation schemas.

Note

The channel field is expected to be an index in derived classes.

class Config

Bases: BaseConfig

name: str | None = 'BaseChannelFeatures'

name of schema

class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for local field potential features.

This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf (Series[float]) – Root mean square of LFP signal in dB.

  • psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).

  • psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).

  • psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).

  • psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).

  • psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).

  • psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).

rms_lf: float = 'rms_lf'
psd_lfp: float = 'psd_lfp'
psd_delta: float = 'psd_delta'
psd_theta: float = 'psd_theta'
psd_alpha: float = 'psd_alpha'
psd_beta: float = 'psd_beta'
psd_gamma: float = 'psd_gamma'
psd_residual_lfp: float | None = 'psd_residual_lfp'
psd_residual_delta: float | None = 'psd_residual_delta'
psd_residual_theta: float | None = 'psd_residual_theta'
psd_residual_alpha: float | None = 'psd_residual_alpha'
psd_residual_beta: float | None = 'psd_residual_beta'
psd_residual_gamma: float | None = 'psd_residual_gamma'
aperiodic_offset: float | None = 'aperiodic_offset'
aperiodic_exponent: float | None = 'aperiodic_exponent'
decay_fit_error: float | None = 'decay_fit_error'
decay_fit_r_squared: float | None = 'decay_fit_r_squared'
decay_n_peaks: float | None = 'decay_n_peaks'
rms_lf_no_car: float | None = 'rms_lf_no_car'
class Config

Bases: Config

name: str | None = 'ModelLfFeatures'

name of schema

class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for current source density features.

This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.

  • psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).

  • psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).

  • psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).

  • psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).

  • psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).

  • psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).

rms_lf_csd: float = 'rms_lf_csd'
psd_delta_csd: float = 'psd_delta_csd'
psd_theta_csd: float = 'psd_theta_csd'
psd_alpha_csd: float = 'psd_alpha_csd'
psd_beta_csd: float = 'psd_beta_csd'
psd_gamma_csd: float = 'psd_gamma_csd'
psd_lfp_csd: float = 'psd_lfp_csd'
rms_lf_csd_diff1: float | None = 'rms_lf_csd_diff1'
psd_delta_csd_diff1: float | None = 'psd_delta_csd_diff1'
psd_theta_csd_diff1: float | None = 'psd_theta_csd_diff1'
psd_alpha_csd_diff1: float | None = 'psd_alpha_csd_diff1'
psd_beta_csd_diff1: float | None = 'psd_beta_csd_diff1'
psd_gamma_csd_diff1: float | None = 'psd_gamma_csd_diff1'
psd_lfp_csd_diff1: float | None = 'psd_lfp_csd_diff1'
class Config

Bases: Config

name: str | None = 'ModelCsdFeatures'

name of schema

class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for action potential features.

This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.

Variables:
  • rms_ap (Series[float]) – Root mean square of AP signal in dB.

  • cor_ratio (Series[float]) – Cross-correlation over auto-correlation ratio.

  • channel_labels (Series[int]) – Quality labels for channels.

rms_ap: float = 'rms_ap'
cor_ratio: float = 'cor_ratio'
channel_labels: int = 'channel_labels'
class Config

Bases: Config

name: str | None = 'ModelApFeatures'

name of schema

class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Spike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).

Variables:
  • alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).

  • alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).

  • depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).

  • polarity (Series[float]) – Spike polarity, dimensionless.

  • recovery_slope (Series[float]) – Slope during recovery phase (V/s).

  • repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).

  • slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).

  • slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).

  • spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).

  • spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).

  • spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).

  • tip_val (Series[float]) – Tip amplitude value (V).

Note

tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.

alpha_mean: float = 'alpha_mean'
alpha_std: float = 'alpha_std'
depolarisation_slope: float = 'depolarisation_slope'
polarity: float = 'polarity'
recovery_slope: float = 'recovery_slope'
repolarisation_slope: float = 'repolarisation_slope'
slowness_s_per_m: float | None = 'slowness_s_per_m'
slowness_s_per_m_std: float | None = 'slowness_s_per_m_std'
spatial_spread_um: float | None = 'spatial_spread_um'
spatial_spread_um_std: float | None = 'spatial_spread_um_std'
spike_count: float = 'spike_count'
tip_val: float = 'tip_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeSharedFeatures'

name of schema

class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for spike waveform features.

This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.

Variables:
  • peak_time_secs (Series[float]) – Time to peak in seconds.

  • peak_val (Series[float]) – Peak amplitude value (V).

  • recovery_time_secs (Series[float]) – Recovery time in seconds.

  • tip_time_secs (Series[float]) – Time to tip in seconds.

  • trough_time_secs (Series[float]) – Time to trough in seconds.

  • trough_val (Series[float]) – Trough amplitude value (V).

Note

peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.

peak_time_secs: float = 'peak_time_secs'
peak_val: float = 'peak_val'
recovery_time_secs: float = 'recovery_time_secs'
tip_time_secs: float = 'tip_time_secs'
trough_time_secs: float = 'trough_time_secs'
trough_val: float = 'trough_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeFeatures'

name of schema

class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for the sparse, neuroscientist-facing spike-shape feature set.

Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.

The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.

Variables:
  • spike_width_secs (Series[float]) – Trough-to-peak spike duration (s).

  • predepolarisation_width_secs (Series[float]) – Tip-to-peak duration (s).

  • spike_amplitude (Series[float]) – Trough-to-peak amplitude (V).

  • peak_to_trough_ratio_log (Series[float]) – Log amplitude ratio (dimensionless).

spike_width_secs: float = 'spike_width_secs'
predepolarisation_width_secs: float = 'predepolarisation_width_secs'
spike_amplitude: float = 'spike_amplitude'
peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
class Config

Bases: Config

name: str | None = 'ModelSpikeShapeFeatures'

name of schema

ephysatlas.features._remap_waveform_shape_arrays(get)[source]

Core of the waveform-shape remapping, agnostic to the container.

Parameters

getcallable

get(name) returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).

Returns

dict

Maps each ModelSpikeShapeFeatures column name to its array.

ephysatlas.features.remap_waveform_shape_features(df)[source]

Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.

The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only recovery_time_secs (exact constant offset of trough_time_secs) and peak_val/trough_val/peak_time_secs/ trough_time_secs/tip_time_secs (reparametrised into spike_amplitude/peak_to_trough_ratio_log/spike_width_secs/ predepolarisation_width_secs without losing relative information) are actually dropped. tip_val and the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information. slowness_s_per_m/ spatial_spread_um (#123) and their _std counterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.

Return type:

DataFrame

Parameters

dfpandas.DataFrame

Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.

Returns

pandas.DataFrame

Columns of ModelSpikeShapeFeatures, same index as df.

ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]

Same transform as remap_waveform_shape_features, for a download_encoding_volume array.

Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.

Parameters

ephys_atlas_volnumpy.ndarray

Shape (nx, ny, nz, n_features).

feature_namesnumpy.ndarray or sequence of str

Length n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.

Returns

remapped_volnumpy.ndarray

Shape (nx, ny, nz, 16), dtype float32.

remapped_feature_namesnumpy.ndarray

Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).

Raises

KeyError

If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).

class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for channel layout information.

This schema defines the structure and validation rules for channel layout features including spatial positioning.

Variables:
  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
class Config

Bases: Config

name: str | None = 'ModelChannelLayout'

name of schema

class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for planned histology coordinates.

This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.

Note

Same coordinate space as ModelHistologyResolved (IBL, metres from Bregma); Populated by feature_computation.add_target_coordinates, which queries Alyx trajectories with provenance__lte,50 and takes the best provenance available per PROVENANCE_PREFERENCE: Micro-manipulator (30), else Planned (10), else Histology track (50); an insertion with none of the three raises ValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.

Variables:
  • x_target (Series[float]) – Target X-coordinate in metres from Bregma (IBL coordinates space).

  • y_target (Series[float]) – Target Y-coordinate in metres from Bregma (IBL coordinates space).

  • z_target (Series[float]) – Target Z-coordinate in metres from Bregma (IBL coordinates space).

x_target: float = 'x_target'
y_target: float = 'y_target'
z_target: float = 'z_target'
class Config

Bases: Config

name: str | None = 'ModelHistologyPlanned'

name of schema

class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for resolved histology coordinates.

This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.

Note

Same coordinate space as ModelHistologyPlanned. These come from SpikeSortingLoader.load_channels() – the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the *_target columns use, hence “at the most recent histology step”.

Variables:
  • x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).

  • atlas_id (Series[int]) – Atlas region identifier.

  • acronym (Series[str]) – Brain region acronym.

x: float = 'x'
y: float = 'y'
z: float = 'z'
atlas_id: int = 'atlas_id'
acronym: str = 'acronym'
class Config

Bases: Config

name: str | None = 'ModelHistologyResolved'

name of schema

class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]

Bases: ModelSpikeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for all raw features.

This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.

Note

This class inherits from multiple feature schemas to provide a unified interface for all feature types.

class Config

Bases: Config

name: str | None = 'ModelRawFeatures'

name of schema

class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]

Bases: ModelSpikeShapeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for denoised features, after the waveform-shape remap.

Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.

Note

Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.

class Config

Bases: Config

name: str | None = 'ModelDenoisedFeatures'

name of schema

class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]

Bases: DataFrameModel

Schema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.

pid: str = 'pid'
eid: str = 'eid'
probe_name: str | None = 'probe_name'
probe_serial: str | None = 'probe_serial'
neuropixel_version: str | None = 'neuropixel_version'
probe_model: str | None = 'probe_model'
referencing_scheme: str | None = 'referencing_scheme'
lab: str | None = 'lab'
pname: str | None = 'pname'
spike_sorting: str | None = 'spike_sorting'
spike_sorting_version: str | None = 'spike_sorting_version'
histology: str | None = 'histology'
record_length: float | None = 'record_length'
channel_count: float | None = 'channel_count'
bwm: bool = 'bwm'
class Config

Bases: BaseConfig

name: str | None = 'ModelProbeDetails'

name of schema

ephysatlas.features.feature_group_of_column_map()[source]

Map each denoisable feature column name to its provenance group.

Return type:

dict

Returns

dict

Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.

ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]

Get list of feature column names by provenance.

This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.

Parameters:

features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.

Returns:

Sorted list of feature column names excluding the ‘channel’ column.

Return type:

list

Note

The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.

ephysatlas.features._get_power_in_band(fscale, period, band)[source]

Calculate power in a specific frequency band.

Parameters:
  • fscale (np.ndarray) – Frequency scale array.

  • period (np.ndarray) – Periodogram values.

  • band (list) – Frequency band [low, high] in Hz.

Returns:

Power in the specified band in dB relative to v/sqrt(Hz).

Return type:

np.ndarray

Note

This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.

ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]

Extract power spectral density decay features from electrophysiological data.

This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.

Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.

The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.

Parameters

datanp.ndarray

Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.

fsfloat

Sampling frequency of the data in Hz.

fscalenp.ndarray

Frequency scale array corresponding to the period data.

periodnp.ndarray

2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.

bandsdict

Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.

npersegint, optional

Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.

PSD_rangelist of float, optional

Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.

Returns

pd.DataFrame

One row per channel. Aperiodic-component columns: aperiodic_offset (log10 power at 1 Hz), aperiodic_exponent (1/f slope in log-log space), decay_fit_error (RMS fit error), decay_fit_r_squared (goodness of fit) and decay_n_peaks (number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal): psd_residual_delta, psd_residual_theta, psd_residual_alpha, psd_residual_beta, psd_residual_gamma and psd_residual_lfp.

Notes

The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.

Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.

The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1

The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))

Examples

>>> import numpy as np
>>> # Generate sample data and frequency arrays
>>> data = np.random.randn(10, 10000)
>>> fs = 1000.0  # 1 kHz sampling rate
>>> fscale = np.linspace(0, 90, 100)
>>> period = np.random.rand(10, 100)  # Mock PSD data
>>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]}
>>> features = get_psd_decay_features(data, fs, fscale, period, bands)
>>> print(features.columns)
Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error',
       'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta',
       'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')

References

Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.

ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]

Compute LF features from a numpy array.

Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,

’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].

Return type:

pd.DataFrame

Note

The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.

ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]

Compute CSD features from a numpy array.

Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

  • decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10. 1 skips scipy.signal.decimate entirely (passed through unchanged) rather than calling it with a no-op factor: scipy.signal.decimate(..., q=1, ...) designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, which scipy.signal.firwin rejects.

  • scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.

  • denoise (bool, optional) – Apply ibldsp.cadzow.cadzow_denoiser to the decimated data before the CSD finite difference. Defaults to True (today’s behavior). Set False when data has already been Cadzow-denoised upstream (e.g. combined with decimate=1 for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,

’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].

Return type:

pd.DataFrame

Note

The function applies Cadzow denoising and current source density computation before computing the spectral features.

ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]

Compute AP features from a numpy array.

Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.

Parameters:
  • data (np.ndarray) – AP band data array with shape (channels, samples).

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • channel_labels (np.ndarray) – Array of channel quality labels.

Returns:

DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].

Return type:

pd.DataFrame

Raises:

AssertionError – If geometry or channel_labels are not provided.

Note

This function computes RMS values and cross-correlation ratios for action potential band data.

ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]

Perform spike detection using Dartsort.

This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.

Parameters:
  • data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • **params – Additional parameters for Dartsort configuration.

Returns:

A tuple containing:

  • df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.

  • d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.

Return type:

tuple

Note

This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.

DARTsort detects spikes on data normalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning: ptp and the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_val and the slopes computed by ibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.

ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]

Dartsort backend for spike detection.

This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Dartsort parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (DartParameters): Dartsort parameters object.

Return type:

tuple

ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]

SpikeInterface backend for spike detection.

This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – SpikeInterface parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (dict): SpikeInterface parameters object.

Return type:

tuple

Raises:

Note

This function is currently a placeholder and needs full implementation for SpikeInterface backend support.

ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]

Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.

Parameters:
  • channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in d_waveforms["channel_index"]. Padded with the sentinel n_channels_real for unused neighbour slots.

  • spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g. df_spikes["channel"].

  • x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.

  • y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.

Returns:

(n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused

neighbour slots.

Return type:

np.ndarray

ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]

Spike detection and feature extraction with multiple backend support.

This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.

  • return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.

  • backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Backend-specific parameters.

Returns:

If return_waveforms is False, returns DataFrame with

aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).

Return type:

pd.DataFrame or tuple

Raises:

ValueError – If an unknown backend is specified.

Note

The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.

ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]

Compute cross-correlation over auto-correlation ratio.

This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.

Parameters:
  • v (np.ndarray) – Voltage array for AP band with shape (nc, ns) where nc is number of channels and ns is number of samples.

  • geometry (dict) – Geometry dictionary with ‘x’ and ‘y’ arrays for electrode positions.

  • n_neighbor (int, optional) – Number of neighboring channels to consider. Defaults to 3.

Returns:

Array of size (nc,) containing the correlation ratios.

Return type:

np.ndarray

Note

The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.

ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]

Denoise AP features using total variation filter.

Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.

Parameters:
  • feature (np.ndarray) – AP feature to denoise with shape (nc,).

  • xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).

  • labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.

  • fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.

Returns:

Denoised AP features with shape (nc,).

Return type:

np.ndarray

Note

This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.

class ephysatlas.features._EphysTransformerInterface[source]

Bases: ABC, OneToOneFeatureMixin, TransformerMixin, BaseEstimator

Abstract base class for electrophysiological feature transformers.

This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.

Note

This is an abstract base class that should not be instantiated directly.

validate_X(X)[source]
Return type:

None

fit_transform(X=None, y=None)[source]

Fit the transformer and transform the data.

Parameters:
  • X (pd.DataFrame, optional) – Input data to fit and transform.

  • y – Ignored, present for compatibility with scikit-learn interface.

Returns:

Transformed data.

Return type:

pd.DataFrame

class ephysatlas.features.EphysTransformer[source]

Bases: _EphysTransformerInterface

fit(X=None, y=None)[source]
transform(X, y=None)[source]
class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Bases: _EphysTransformerInterface

__init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

TV-denoise electrophysiological features, with one weight factor per feature group.

Parameters

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.

channel_labelsnp.ndarray, optional

Channel quality annotation array with shape (nc,).

transform(X, y=None)[source]
fit(X=None, y=None)[source]
ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Applies total variation filter denoising to the features of a single probe insertion dataframe.

This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.

Parameters

df_pidpandas.DataFrame

DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.

Returns

pandas.DataFrame

A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.

Transformers

Electrophysiological feature extraction and processing module.

This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.

The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility

Classes

DartParameters

Configuration parameters for Dartsort backend

ChannelDataFrameSchema

Pandera schema for channel data validation

ModelLfFeatures

Schema for local field potential features

ModelCsdFeatures

Schema for current source density features

ModelApFeatures

Schema for action potential features

ModelSpikeFeatures

Schema for spike waveform features

ModelSpikeShapeFeatures

Sparse, neuroscientist-facing remapping of ModelSpikeFeatures

ModelChannelLayout

Schema for channel layout information

ModelHistologyPlanned

Schema for planned histology coordinates

ModelHistologyResolved

Schema for resolved histology coordinates

ModelRawFeatures

Combined schema for all raw features

EphysTransformer

Transformer for applying feature transformations

EphysDenoiser

Transformer for denoising electrophysiological features

Functions

_setup_scratch_directory

Set up scratch directory with fallback logic

voltage_features_set

Get list of feature column names by provenance

lf

Compute LF features from numpy array

csd

Compute CSD features from numpy array

ap

Compute AP features from numpy array

dart_subtraction_numpy

Perform spike detection using Dartsort

spikes

Spike detection and feature extraction with multiple backend support

remap_waveform_shape_features

Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set

remap_waveform_shape_features_volume

Same remapping, for a brainwide encoding volume instead of a dataframe

xcor_acor_ratio

Compute cross-correlation over auto-correlation ratio

denoise_shank

Denoise AP features using total variation filter

denoise_dataframe

Apply total variation filter denoising to features

Constants

__features_version__str

Version of the feature extractor code

BANDSdict

Frequency bands for spectral analysis

FEATURES_LISTlist

List of available feature types

Examples

>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000)  # 64 channels, 30k samples
>>> fs = 30000  # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))

Notes

This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.

See Also

ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities

ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]

Set up scratch directory with fallback logic.

Parameters:

scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.

Returns:

Path to the created scratch directory.

Return type:

Path

Note

This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.

ephysatlas.features.get_feature_cmin(feature_name)[source]

Get the minimum value for a given feature.

Parameters:

feature_name (str) – Name of the feature.

Returns:

Minimum value for the feature.

Return type:

float

Note

This function is currently a placeholder and needs implementation.

class ephysatlas.features.DartParameters(**data)[source]

Bases: BaseModel

Configuration parameters for Dartsort backend.

This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.

Variables:
  • localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.

  • chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).

  • trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.

  • scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.

  • n_jobs (int) – DARTsort worker count used to build its ComputationConfig. 0 runs in the main process; positive values select that many workers. Defaults to 0.

  • detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as voltage_threshold. Defaults to 4.0.

  • spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.

  • positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.

  • residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as subtraction_threshold. Defaults to 3.162 (sqrt(10)).

localization_radius: Annotated[float, Gt(gt=0)]
chunk_length_samples: Annotated[int, Gt(gt=0)]
trough_offset: Annotated[int, Gt(gt=0)]
detection_threshold: Annotated[float, Gt(gt=0)]
spatial_dedup_radius_um: Annotated[float, Gt(gt=0)]
positive_temporal_dedup_radius_samples: Annotated[int, Gt(gt=0)]
residnorm_decrease_threshold: Annotated[float, Gt(gt=0)]
scratch_dir: Path | str | None
n_jobs: Annotated[int, Ge(ge=0)]
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]

Bases: DataFrameModel

Pandera schema for channel data validation.

This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.

Note

x/y/z are in metres, while axial_um/lateral_um are in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.

Variables:
  • pid (Series[str]) – Probe insertion ID.

  • channel (Series[int]) – Channel index.

  • x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).

  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

  • acronym (Series[str]) – Brain region acronym.

  • atlas_id (Series[int]) – Atlas region identifier.

pid: str = 'pid'
channel: int = 'channel'
x: float = 'x'
y: float = 'y'
z: float = 'z'
axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
acronym: str = 'acronym'
atlas_id: int = 'atlas_id'
distance_to_tip_um: float | None = 'distance_to_tip_um'
class Config

Bases: BaseConfig

name: str | None = 'ChannelDataFrameSchema'

name of schema

class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]

Bases: DataFrameModel

Base class for channel-based feature schemas.

This is an abstract base class that provides the foundation for all channel-based feature validation schemas.

Note

The channel field is expected to be an index in derived classes.

class Config

Bases: BaseConfig

name: str | None = 'BaseChannelFeatures'

name of schema

class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for local field potential features.

This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf (Series[float]) – Root mean square of LFP signal in dB.

  • psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).

  • psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).

  • psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).

  • psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).

  • psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).

  • psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).

rms_lf: float = 'rms_lf'
psd_lfp: float = 'psd_lfp'
psd_delta: float = 'psd_delta'
psd_theta: float = 'psd_theta'
psd_alpha: float = 'psd_alpha'
psd_beta: float = 'psd_beta'
psd_gamma: float = 'psd_gamma'
psd_residual_lfp: float | None = 'psd_residual_lfp'
psd_residual_delta: float | None = 'psd_residual_delta'
psd_residual_theta: float | None = 'psd_residual_theta'
psd_residual_alpha: float | None = 'psd_residual_alpha'
psd_residual_beta: float | None = 'psd_residual_beta'
psd_residual_gamma: float | None = 'psd_residual_gamma'
aperiodic_offset: float | None = 'aperiodic_offset'
aperiodic_exponent: float | None = 'aperiodic_exponent'
decay_fit_error: float | None = 'decay_fit_error'
decay_fit_r_squared: float | None = 'decay_fit_r_squared'
decay_n_peaks: float | None = 'decay_n_peaks'
rms_lf_no_car: float | None = 'rms_lf_no_car'
class Config

Bases: Config

name: str | None = 'ModelLfFeatures'

name of schema

class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for current source density features.

This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.

  • psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).

  • psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).

  • psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).

  • psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).

  • psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).

  • psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).

rms_lf_csd: float = 'rms_lf_csd'
psd_delta_csd: float = 'psd_delta_csd'
psd_theta_csd: float = 'psd_theta_csd'
psd_alpha_csd: float = 'psd_alpha_csd'
psd_beta_csd: float = 'psd_beta_csd'
psd_gamma_csd: float = 'psd_gamma_csd'
psd_lfp_csd: float = 'psd_lfp_csd'
rms_lf_csd_diff1: float | None = 'rms_lf_csd_diff1'
psd_delta_csd_diff1: float | None = 'psd_delta_csd_diff1'
psd_theta_csd_diff1: float | None = 'psd_theta_csd_diff1'
psd_alpha_csd_diff1: float | None = 'psd_alpha_csd_diff1'
psd_beta_csd_diff1: float | None = 'psd_beta_csd_diff1'
psd_gamma_csd_diff1: float | None = 'psd_gamma_csd_diff1'
psd_lfp_csd_diff1: float | None = 'psd_lfp_csd_diff1'
class Config

Bases: Config

name: str | None = 'ModelCsdFeatures'

name of schema

class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for action potential features.

This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.

Variables:
  • rms_ap (Series[float]) – Root mean square of AP signal in dB.

  • cor_ratio (Series[float]) – Cross-correlation over auto-correlation ratio.

  • channel_labels (Series[int]) – Quality labels for channels.

rms_ap: float = 'rms_ap'
cor_ratio: float = 'cor_ratio'
channel_labels: int = 'channel_labels'
class Config

Bases: Config

name: str | None = 'ModelApFeatures'

name of schema

class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Spike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).

Variables:
  • alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).

  • alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).

  • depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).

  • polarity (Series[float]) – Spike polarity, dimensionless.

  • recovery_slope (Series[float]) – Slope during recovery phase (V/s).

  • repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).

  • slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).

  • slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).

  • spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).

  • spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).

  • spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).

  • tip_val (Series[float]) – Tip amplitude value (V).

Note

tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.

alpha_mean: float = 'alpha_mean'
alpha_std: float = 'alpha_std'
depolarisation_slope: float = 'depolarisation_slope'
polarity: float = 'polarity'
recovery_slope: float = 'recovery_slope'
repolarisation_slope: float = 'repolarisation_slope'
slowness_s_per_m: float | None = 'slowness_s_per_m'
slowness_s_per_m_std: float | None = 'slowness_s_per_m_std'
spatial_spread_um: float | None = 'spatial_spread_um'
spatial_spread_um_std: float | None = 'spatial_spread_um_std'
spike_count: float = 'spike_count'
tip_val: float = 'tip_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeSharedFeatures'

name of schema

class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for spike waveform features.

This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.

Variables:
  • peak_time_secs (Series[float]) – Time to peak in seconds.

  • peak_val (Series[float]) – Peak amplitude value (V).

  • recovery_time_secs (Series[float]) – Recovery time in seconds.

  • tip_time_secs (Series[float]) – Time to tip in seconds.

  • trough_time_secs (Series[float]) – Time to trough in seconds.

  • trough_val (Series[float]) – Trough amplitude value (V).

Note

peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.

peak_time_secs: float = 'peak_time_secs'
peak_val: float = 'peak_val'
recovery_time_secs: float = 'recovery_time_secs'
tip_time_secs: float = 'tip_time_secs'
trough_time_secs: float = 'trough_time_secs'
trough_val: float = 'trough_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeFeatures'

name of schema

class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for the sparse, neuroscientist-facing spike-shape feature set.

Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.

The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.

Variables:
  • spike_width_secs (Series[float]) – Trough-to-peak spike duration (s).

  • predepolarisation_width_secs (Series[float]) – Tip-to-peak duration (s).

  • spike_amplitude (Series[float]) – Trough-to-peak amplitude (V).

  • peak_to_trough_ratio_log (Series[float]) – Log amplitude ratio (dimensionless).

spike_width_secs: float = 'spike_width_secs'
predepolarisation_width_secs: float = 'predepolarisation_width_secs'
spike_amplitude: float = 'spike_amplitude'
peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
class Config

Bases: Config

name: str | None = 'ModelSpikeShapeFeatures'

name of schema

ephysatlas.features._remap_waveform_shape_arrays(get)[source]

Core of the waveform-shape remapping, agnostic to the container.

Parameters

getcallable

get(name) returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).

Returns

dict

Maps each ModelSpikeShapeFeatures column name to its array.

ephysatlas.features.remap_waveform_shape_features(df)[source]

Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.

The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only recovery_time_secs (exact constant offset of trough_time_secs) and peak_val/trough_val/peak_time_secs/ trough_time_secs/tip_time_secs (reparametrised into spike_amplitude/peak_to_trough_ratio_log/spike_width_secs/ predepolarisation_width_secs without losing relative information) are actually dropped. tip_val and the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information. slowness_s_per_m/ spatial_spread_um (#123) and their _std counterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.

Return type:

DataFrame

Parameters

dfpandas.DataFrame

Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.

Returns

pandas.DataFrame

Columns of ModelSpikeShapeFeatures, same index as df.

ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]

Same transform as remap_waveform_shape_features, for a download_encoding_volume array.

Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.

Parameters

ephys_atlas_volnumpy.ndarray

Shape (nx, ny, nz, n_features).

feature_namesnumpy.ndarray or sequence of str

Length n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.

Returns

remapped_volnumpy.ndarray

Shape (nx, ny, nz, 16), dtype float32.

remapped_feature_namesnumpy.ndarray

Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).

Raises

KeyError

If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).

class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for channel layout information.

This schema defines the structure and validation rules for channel layout features including spatial positioning.

Variables:
  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
class Config

Bases: Config

name: str | None = 'ModelChannelLayout'

name of schema

class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for planned histology coordinates.

This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.

Note

Same coordinate space as ModelHistologyResolved (IBL, metres from Bregma); Populated by feature_computation.add_target_coordinates, which queries Alyx trajectories with provenance__lte,50 and takes the best provenance available per PROVENANCE_PREFERENCE: Micro-manipulator (30), else Planned (10), else Histology track (50); an insertion with none of the three raises ValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.

Variables:
  • x_target (Series[float]) – Target X-coordinate in metres from Bregma (IBL coordinates space).

  • y_target (Series[float]) – Target Y-coordinate in metres from Bregma (IBL coordinates space).

  • z_target (Series[float]) – Target Z-coordinate in metres from Bregma (IBL coordinates space).

x_target: float = 'x_target'
y_target: float = 'y_target'
z_target: float = 'z_target'
class Config

Bases: Config

name: str | None = 'ModelHistologyPlanned'

name of schema

class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for resolved histology coordinates.

This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.

Note

Same coordinate space as ModelHistologyPlanned. These come from SpikeSortingLoader.load_channels() – the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the *_target columns use, hence “at the most recent histology step”.

Variables:
  • x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).

  • atlas_id (Series[int]) – Atlas region identifier.

  • acronym (Series[str]) – Brain region acronym.

x: float = 'x'
y: float = 'y'
z: float = 'z'
atlas_id: int = 'atlas_id'
acronym: str = 'acronym'
class Config

Bases: Config

name: str | None = 'ModelHistologyResolved'

name of schema

class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]

Bases: ModelSpikeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for all raw features.

This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.

Note

This class inherits from multiple feature schemas to provide a unified interface for all feature types.

class Config

Bases: Config

name: str | None = 'ModelRawFeatures'

name of schema

class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]

Bases: ModelSpikeShapeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for denoised features, after the waveform-shape remap.

Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.

Note

Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.

class Config

Bases: Config

name: str | None = 'ModelDenoisedFeatures'

name of schema

class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]

Bases: DataFrameModel

Schema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.

pid: str = 'pid'
eid: str = 'eid'
probe_name: str | None = 'probe_name'
probe_serial: str | None = 'probe_serial'
neuropixel_version: str | None = 'neuropixel_version'
probe_model: str | None = 'probe_model'
referencing_scheme: str | None = 'referencing_scheme'
lab: str | None = 'lab'
pname: str | None = 'pname'
spike_sorting: str | None = 'spike_sorting'
spike_sorting_version: str | None = 'spike_sorting_version'
histology: str | None = 'histology'
record_length: float | None = 'record_length'
channel_count: float | None = 'channel_count'
bwm: bool = 'bwm'
class Config

Bases: BaseConfig

name: str | None = 'ModelProbeDetails'

name of schema

ephysatlas.features.feature_group_of_column_map()[source]

Map each denoisable feature column name to its provenance group.

Return type:

dict

Returns

dict

Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.

ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]

Get list of feature column names by provenance.

This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.

Parameters:

features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.

Returns:

Sorted list of feature column names excluding the ‘channel’ column.

Return type:

list

Note

The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.

ephysatlas.features._get_power_in_band(fscale, period, band)[source]

Calculate power in a specific frequency band.

Parameters:
  • fscale (np.ndarray) – Frequency scale array.

  • period (np.ndarray) – Periodogram values.

  • band (list) – Frequency band [low, high] in Hz.

Returns:

Power in the specified band in dB relative to v/sqrt(Hz).

Return type:

np.ndarray

Note

This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.

ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]

Extract power spectral density decay features from electrophysiological data.

This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.

Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.

The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.

Parameters

datanp.ndarray

Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.

fsfloat

Sampling frequency of the data in Hz.

fscalenp.ndarray

Frequency scale array corresponding to the period data.

periodnp.ndarray

2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.

bandsdict

Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.

npersegint, optional

Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.

PSD_rangelist of float, optional

Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.

Returns

pd.DataFrame

One row per channel. Aperiodic-component columns: aperiodic_offset (log10 power at 1 Hz), aperiodic_exponent (1/f slope in log-log space), decay_fit_error (RMS fit error), decay_fit_r_squared (goodness of fit) and decay_n_peaks (number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal): psd_residual_delta, psd_residual_theta, psd_residual_alpha, psd_residual_beta, psd_residual_gamma and psd_residual_lfp.

Notes

The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.

Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.

The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1

The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))

Examples

>>> import numpy as np
>>> # Generate sample data and frequency arrays
>>> data = np.random.randn(10, 10000)
>>> fs = 1000.0  # 1 kHz sampling rate
>>> fscale = np.linspace(0, 90, 100)
>>> period = np.random.rand(10, 100)  # Mock PSD data
>>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]}
>>> features = get_psd_decay_features(data, fs, fscale, period, bands)
>>> print(features.columns)
Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error',
       'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta',
       'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')

References

Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.

ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]

Compute LF features from a numpy array.

Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,

’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].

Return type:

pd.DataFrame

Note

The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.

ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]

Compute CSD features from a numpy array.

Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

  • decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10. 1 skips scipy.signal.decimate entirely (passed through unchanged) rather than calling it with a no-op factor: scipy.signal.decimate(..., q=1, ...) designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, which scipy.signal.firwin rejects.

  • scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.

  • denoise (bool, optional) – Apply ibldsp.cadzow.cadzow_denoiser to the decimated data before the CSD finite difference. Defaults to True (today’s behavior). Set False when data has already been Cadzow-denoised upstream (e.g. combined with decimate=1 for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,

’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].

Return type:

pd.DataFrame

Note

The function applies Cadzow denoising and current source density computation before computing the spectral features.

ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]

Compute AP features from a numpy array.

Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.

Parameters:
  • data (np.ndarray) – AP band data array with shape (channels, samples).

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • channel_labels (np.ndarray) – Array of channel quality labels.

Returns:

DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].

Return type:

pd.DataFrame

Raises:

AssertionError – If geometry or channel_labels are not provided.

Note

This function computes RMS values and cross-correlation ratios for action potential band data.

ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]

Perform spike detection using Dartsort.

This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.

Parameters:
  • data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • **params – Additional parameters for Dartsort configuration.

Returns:

A tuple containing:

  • df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.

  • d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.

Return type:

tuple

Note

This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.

DARTsort detects spikes on data normalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning: ptp and the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_val and the slopes computed by ibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.

ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]

Dartsort backend for spike detection.

This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Dartsort parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (DartParameters): Dartsort parameters object.

Return type:

tuple

ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]

SpikeInterface backend for spike detection.

This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – SpikeInterface parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (dict): SpikeInterface parameters object.

Return type:

tuple

Raises:

Note

This function is currently a placeholder and needs full implementation for SpikeInterface backend support.

ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]

Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.

Parameters:
  • channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in d_waveforms["channel_index"]. Padded with the sentinel n_channels_real for unused neighbour slots.

  • spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g. df_spikes["channel"].

  • x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.

  • y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.

Returns:

(n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused

neighbour slots.

Return type:

np.ndarray

ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]

Spike detection and feature extraction with multiple backend support.

This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.

  • return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.

  • backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Backend-specific parameters.

Returns:

If return_waveforms is False, returns DataFrame with

aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).

Return type:

pd.DataFrame or tuple

Raises:

ValueError – If an unknown backend is specified.

Note

The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.

ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]

Compute cross-correlation over auto-correlation ratio.

This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.

Parameters:
  • v (np.ndarray) – Voltage array for AP band with shape (nc, ns) where nc is number of channels and ns is number of samples.

  • geometry (dict) – Geometry dictionary with ‘x’ and ‘y’ arrays for electrode positions.

  • n_neighbor (int, optional) – Number of neighboring channels to consider. Defaults to 3.

Returns:

Array of size (nc,) containing the correlation ratios.

Return type:

np.ndarray

Note

The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.

ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]

Denoise AP features using total variation filter.

Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.

Parameters:
  • feature (np.ndarray) – AP feature to denoise with shape (nc,).

  • xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).

  • labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.

  • fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.

Returns:

Denoised AP features with shape (nc,).

Return type:

np.ndarray

Note

This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.

class ephysatlas.features._EphysTransformerInterface[source]

Bases: ABC, OneToOneFeatureMixin, TransformerMixin, BaseEstimator

Abstract base class for electrophysiological feature transformers.

This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.

Note

This is an abstract base class that should not be instantiated directly.

validate_X(X)[source]
Return type:

None

fit_transform(X=None, y=None)[source]

Fit the transformer and transform the data.

Parameters:
  • X (pd.DataFrame, optional) – Input data to fit and transform.

  • y – Ignored, present for compatibility with scikit-learn interface.

Returns:

Transformed data.

Return type:

pd.DataFrame

class ephysatlas.features.EphysTransformer[source]

Bases: _EphysTransformerInterface

fit(X=None, y=None)[source]
transform(X, y=None)[source]
class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Bases: _EphysTransformerInterface

__init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

TV-denoise electrophysiological features, with one weight factor per feature group.

Parameters

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.

channel_labelsnp.ndarray, optional

Channel quality annotation array with shape (nc,).

transform(X, y=None)[source]
fit(X=None, y=None)[source]
ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Applies total variation filter denoising to the features of a single probe insertion dataframe.

This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.

Parameters

df_pidpandas.DataFrame

DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.

Returns

pandas.DataFrame

A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.

Configuration

Electrophysiological feature extraction and processing module.

This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.

The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility

Classes

DartParameters

Configuration parameters for Dartsort backend

ChannelDataFrameSchema

Pandera schema for channel data validation

ModelLfFeatures

Schema for local field potential features

ModelCsdFeatures

Schema for current source density features

ModelApFeatures

Schema for action potential features

ModelSpikeFeatures

Schema for spike waveform features

ModelSpikeShapeFeatures

Sparse, neuroscientist-facing remapping of ModelSpikeFeatures

ModelChannelLayout

Schema for channel layout information

ModelHistologyPlanned

Schema for planned histology coordinates

ModelHistologyResolved

Schema for resolved histology coordinates

ModelRawFeatures

Combined schema for all raw features

EphysTransformer

Transformer for applying feature transformations

EphysDenoiser

Transformer for denoising electrophysiological features

Functions

_setup_scratch_directory

Set up scratch directory with fallback logic

voltage_features_set

Get list of feature column names by provenance

lf

Compute LF features from numpy array

csd

Compute CSD features from numpy array

ap

Compute AP features from numpy array

dart_subtraction_numpy

Perform spike detection using Dartsort

spikes

Spike detection and feature extraction with multiple backend support

remap_waveform_shape_features

Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set

remap_waveform_shape_features_volume

Same remapping, for a brainwide encoding volume instead of a dataframe

xcor_acor_ratio

Compute cross-correlation over auto-correlation ratio

denoise_shank

Denoise AP features using total variation filter

denoise_dataframe

Apply total variation filter denoising to features

Constants

__features_version__str

Version of the feature extractor code

BANDSdict

Frequency bands for spectral analysis

FEATURES_LISTlist

List of available feature types

Examples

>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000)  # 64 channels, 30k samples
>>> fs = 30000  # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))

Notes

This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.

See Also

ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities

ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]

Set up scratch directory with fallback logic.

Parameters:

scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.

Returns:

Path to the created scratch directory.

Return type:

Path

Note

This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.

ephysatlas.features.get_feature_cmin(feature_name)[source]

Get the minimum value for a given feature.

Parameters:

feature_name (str) – Name of the feature.

Returns:

Minimum value for the feature.

Return type:

float

Note

This function is currently a placeholder and needs implementation.

class ephysatlas.features.DartParameters(**data)[source]

Bases: BaseModel

Configuration parameters for Dartsort backend.

This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.

Variables:
  • localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.

  • chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).

  • trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.

  • scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.

  • n_jobs (int) – DARTsort worker count used to build its ComputationConfig. 0 runs in the main process; positive values select that many workers. Defaults to 0.

  • detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as voltage_threshold. Defaults to 4.0.

  • spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.

  • positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.

  • residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as subtraction_threshold. Defaults to 3.162 (sqrt(10)).

localization_radius: Annotated[float, Gt(gt=0)]
chunk_length_samples: Annotated[int, Gt(gt=0)]
trough_offset: Annotated[int, Gt(gt=0)]
detection_threshold: Annotated[float, Gt(gt=0)]
spatial_dedup_radius_um: Annotated[float, Gt(gt=0)]
positive_temporal_dedup_radius_samples: Annotated[int, Gt(gt=0)]
residnorm_decrease_threshold: Annotated[float, Gt(gt=0)]
scratch_dir: Path | str | None
n_jobs: Annotated[int, Ge(ge=0)]
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]

Bases: DataFrameModel

Pandera schema for channel data validation.

This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.

Note

x/y/z are in metres, while axial_um/lateral_um are in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.

Variables:
  • pid (Series[str]) – Probe insertion ID.

  • channel (Series[int]) – Channel index.

  • x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).

  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

  • acronym (Series[str]) – Brain region acronym.

  • atlas_id (Series[int]) – Atlas region identifier.

pid: str = 'pid'
channel: int = 'channel'
x: float = 'x'
y: float = 'y'
z: float = 'z'
axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
acronym: str = 'acronym'
atlas_id: int = 'atlas_id'
distance_to_tip_um: float | None = 'distance_to_tip_um'
class Config

Bases: BaseConfig

name: str | None = 'ChannelDataFrameSchema'

name of schema

class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]

Bases: DataFrameModel

Base class for channel-based feature schemas.

This is an abstract base class that provides the foundation for all channel-based feature validation schemas.

Note

The channel field is expected to be an index in derived classes.

class Config

Bases: BaseConfig

name: str | None = 'BaseChannelFeatures'

name of schema

class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for local field potential features.

This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf (Series[float]) – Root mean square of LFP signal in dB.

  • psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).

  • psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).

  • psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).

  • psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).

  • psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).

  • psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).

rms_lf: float = 'rms_lf'
psd_lfp: float = 'psd_lfp'
psd_delta: float = 'psd_delta'
psd_theta: float = 'psd_theta'
psd_alpha: float = 'psd_alpha'
psd_beta: float = 'psd_beta'
psd_gamma: float = 'psd_gamma'
psd_residual_lfp: float | None = 'psd_residual_lfp'
psd_residual_delta: float | None = 'psd_residual_delta'
psd_residual_theta: float | None = 'psd_residual_theta'
psd_residual_alpha: float | None = 'psd_residual_alpha'
psd_residual_beta: float | None = 'psd_residual_beta'
psd_residual_gamma: float | None = 'psd_residual_gamma'
aperiodic_offset: float | None = 'aperiodic_offset'
aperiodic_exponent: float | None = 'aperiodic_exponent'
decay_fit_error: float | None = 'decay_fit_error'
decay_fit_r_squared: float | None = 'decay_fit_r_squared'
decay_n_peaks: float | None = 'decay_n_peaks'
rms_lf_no_car: float | None = 'rms_lf_no_car'
class Config

Bases: Config

name: str | None = 'ModelLfFeatures'

name of schema

class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for current source density features.

This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.

Variables:
  • rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.

  • psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).

  • psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).

  • psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).

  • psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).

  • psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).

  • psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).

rms_lf_csd: float = 'rms_lf_csd'
psd_delta_csd: float = 'psd_delta_csd'
psd_theta_csd: float = 'psd_theta_csd'
psd_alpha_csd: float = 'psd_alpha_csd'
psd_beta_csd: float = 'psd_beta_csd'
psd_gamma_csd: float = 'psd_gamma_csd'
psd_lfp_csd: float = 'psd_lfp_csd'
rms_lf_csd_diff1: float | None = 'rms_lf_csd_diff1'
psd_delta_csd_diff1: float | None = 'psd_delta_csd_diff1'
psd_theta_csd_diff1: float | None = 'psd_theta_csd_diff1'
psd_alpha_csd_diff1: float | None = 'psd_alpha_csd_diff1'
psd_beta_csd_diff1: float | None = 'psd_beta_csd_diff1'
psd_gamma_csd_diff1: float | None = 'psd_gamma_csd_diff1'
psd_lfp_csd_diff1: float | None = 'psd_lfp_csd_diff1'
class Config

Bases: Config

name: str | None = 'ModelCsdFeatures'

name of schema

class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for action potential features.

This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.

Variables:
  • rms_ap (Series[float]) – Root mean square of AP signal in dB.

  • cor_ratio (Series[float]) – Cross-correlation over auto-correlation ratio.

  • channel_labels (Series[int]) – Quality labels for channels.

rms_ap: float = 'rms_ap'
cor_ratio: float = 'cor_ratio'
channel_labels: int = 'channel_labels'
class Config

Bases: Config

name: str | None = 'ModelApFeatures'

name of schema

class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Spike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).

Variables:
  • alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).

  • alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).

  • depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).

  • polarity (Series[float]) – Spike polarity, dimensionless.

  • recovery_slope (Series[float]) – Slope during recovery phase (V/s).

  • repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).

  • slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).

  • slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).

  • spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).

  • spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).

  • spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).

  • tip_val (Series[float]) – Tip amplitude value (V).

Note

tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.

alpha_mean: float = 'alpha_mean'
alpha_std: float = 'alpha_std'
depolarisation_slope: float = 'depolarisation_slope'
polarity: float = 'polarity'
recovery_slope: float = 'recovery_slope'
repolarisation_slope: float = 'repolarisation_slope'
slowness_s_per_m: float | None = 'slowness_s_per_m'
slowness_s_per_m_std: float | None = 'slowness_s_per_m_std'
spatial_spread_um: float | None = 'spatial_spread_um'
spatial_spread_um_std: float | None = 'spatial_spread_um_std'
spike_count: float = 'spike_count'
tip_val: float = 'tip_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeSharedFeatures'

name of schema

class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for spike waveform features.

This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.

Variables:
  • peak_time_secs (Series[float]) – Time to peak in seconds.

  • peak_val (Series[float]) – Peak amplitude value (V).

  • recovery_time_secs (Series[float]) – Recovery time in seconds.

  • tip_time_secs (Series[float]) – Time to tip in seconds.

  • trough_time_secs (Series[float]) – Time to trough in seconds.

  • trough_val (Series[float]) – Trough amplitude value (V).

Note

peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.

peak_time_secs: float = 'peak_time_secs'
peak_val: float = 'peak_val'
recovery_time_secs: float = 'recovery_time_secs'
tip_time_secs: float = 'tip_time_secs'
trough_time_secs: float = 'trough_time_secs'
trough_val: float = 'trough_val'
class Config

Bases: Config

name: str | None = 'ModelSpikeFeatures'

name of schema

class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]

Bases: ModelSpikeSharedFeatures

Schema for the sparse, neuroscientist-facing spike-shape feature set.

Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.

The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.

Variables:
  • spike_width_secs (Series[float]) – Trough-to-peak spike duration (s).

  • predepolarisation_width_secs (Series[float]) – Tip-to-peak duration (s).

  • spike_amplitude (Series[float]) – Trough-to-peak amplitude (V).

  • peak_to_trough_ratio_log (Series[float]) – Log amplitude ratio (dimensionless).

spike_width_secs: float = 'spike_width_secs'
predepolarisation_width_secs: float = 'predepolarisation_width_secs'
spike_amplitude: float = 'spike_amplitude'
peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
class Config

Bases: Config

name: str | None = 'ModelSpikeShapeFeatures'

name of schema

ephysatlas.features._remap_waveform_shape_arrays(get)[source]

Core of the waveform-shape remapping, agnostic to the container.

Parameters

getcallable

get(name) returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).

Returns

dict

Maps each ModelSpikeShapeFeatures column name to its array.

ephysatlas.features.remap_waveform_shape_features(df)[source]

Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.

The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only recovery_time_secs (exact constant offset of trough_time_secs) and peak_val/trough_val/peak_time_secs/ trough_time_secs/tip_time_secs (reparametrised into spike_amplitude/peak_to_trough_ratio_log/spike_width_secs/ predepolarisation_width_secs without losing relative information) are actually dropped. tip_val and the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information. slowness_s_per_m/ spatial_spread_um (#123) and their _std counterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.

Return type:

DataFrame

Parameters

dfpandas.DataFrame

Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.

Returns

pandas.DataFrame

Columns of ModelSpikeShapeFeatures, same index as df.

ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]

Same transform as remap_waveform_shape_features, for a download_encoding_volume array.

Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.

Parameters

ephys_atlas_volnumpy.ndarray

Shape (nx, ny, nz, n_features).

feature_namesnumpy.ndarray or sequence of str

Length n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.

Returns

remapped_volnumpy.ndarray

Shape (nx, ny, nz, 16), dtype float32.

remapped_feature_namesnumpy.ndarray

Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).

Raises

KeyError

If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).

class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for channel layout information.

This schema defines the structure and validation rules for channel layout features including spatial positioning.

Variables:
  • axial_um (Series[float]) – Axial distance in micrometers.

  • lateral_um (Series[float]) – Lateral distance in micrometers.

axial_um: float = 'axial_um'
lateral_um: float = 'lateral_um'
class Config

Bases: Config

name: str | None = 'ModelChannelLayout'

name of schema

class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for planned histology coordinates.

This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.

Note

Same coordinate space as ModelHistologyResolved (IBL, metres from Bregma); Populated by feature_computation.add_target_coordinates, which queries Alyx trajectories with provenance__lte,50 and takes the best provenance available per PROVENANCE_PREFERENCE: Micro-manipulator (30), else Planned (10), else Histology track (50); an insertion with none of the three raises ValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.

Variables:
  • x_target (Series[float]) – Target X-coordinate in metres from Bregma (IBL coordinates space).

  • y_target (Series[float]) – Target Y-coordinate in metres from Bregma (IBL coordinates space).

  • z_target (Series[float]) – Target Z-coordinate in metres from Bregma (IBL coordinates space).

x_target: float = 'x_target'
y_target: float = 'y_target'
z_target: float = 'z_target'
class Config

Bases: Config

name: str | None = 'ModelHistologyPlanned'

name of schema

class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]

Bases: BaseChannelFeatures

Schema for resolved histology coordinates.

This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.

Note

Same coordinate space as ModelHistologyPlanned. These come from SpikeSortingLoader.load_channels() – the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the *_target columns use, hence “at the most recent histology step”.

Variables:
  • x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).

  • y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).

  • z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).

  • atlas_id (Series[int]) – Atlas region identifier.

  • acronym (Series[str]) – Brain region acronym.

x: float = 'x'
y: float = 'y'
z: float = 'z'
atlas_id: int = 'atlas_id'
acronym: str = 'acronym'
class Config

Bases: Config

name: str | None = 'ModelHistologyResolved'

name of schema

class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]

Bases: ModelSpikeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for all raw features.

This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.

Note

This class inherits from multiple feature schemas to provide a unified interface for all feature types.

class Config

Bases: Config

name: str | None = 'ModelRawFeatures'

name of schema

class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]

Bases: ModelSpikeShapeFeatures, ModelCsdFeatures, ModelApFeatures, ModelLfFeatures, ModelChannelLayout

Combined schema for denoised features, after the waveform-shape remap.

Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.

Note

Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.

class Config

Bases: Config

name: str | None = 'ModelDenoisedFeatures'

name of schema

class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]

Bases: DataFrameModel

Schema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.

pid: str = 'pid'
eid: str = 'eid'
probe_name: str | None = 'probe_name'
probe_serial: str | None = 'probe_serial'
neuropixel_version: str | None = 'neuropixel_version'
probe_model: str | None = 'probe_model'
referencing_scheme: str | None = 'referencing_scheme'
lab: str | None = 'lab'
pname: str | None = 'pname'
spike_sorting: str | None = 'spike_sorting'
spike_sorting_version: str | None = 'spike_sorting_version'
histology: str | None = 'histology'
record_length: float | None = 'record_length'
channel_count: float | None = 'channel_count'
bwm: bool = 'bwm'
class Config

Bases: BaseConfig

name: str | None = 'ModelProbeDetails'

name of schema

ephysatlas.features.feature_group_of_column_map()[source]

Map each denoisable feature column name to its provenance group.

Return type:

dict

Returns

dict

Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.

ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]

Get list of feature column names by provenance.

This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.

Parameters:

features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.

Returns:

Sorted list of feature column names excluding the ‘channel’ column.

Return type:

list

Note

The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.

ephysatlas.features._get_power_in_band(fscale, period, band)[source]

Calculate power in a specific frequency band.

Parameters:
  • fscale (np.ndarray) – Frequency scale array.

  • period (np.ndarray) – Periodogram values.

  • band (list) – Frequency band [low, high] in Hz.

Returns:

Power in the specified band in dB relative to v/sqrt(Hz).

Return type:

np.ndarray

Note

This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.

ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]

Extract power spectral density decay features from electrophysiological data.

This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.

Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.

The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.

Parameters

datanp.ndarray

Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.

fsfloat

Sampling frequency of the data in Hz.

fscalenp.ndarray

Frequency scale array corresponding to the period data.

periodnp.ndarray

2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.

bandsdict

Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.

npersegint, optional

Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.

PSD_rangelist of float, optional

Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.

Returns

pd.DataFrame

One row per channel. Aperiodic-component columns: aperiodic_offset (log10 power at 1 Hz), aperiodic_exponent (1/f slope in log-log space), decay_fit_error (RMS fit error), decay_fit_r_squared (goodness of fit) and decay_n_peaks (number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal): psd_residual_delta, psd_residual_theta, psd_residual_alpha, psd_residual_beta, psd_residual_gamma and psd_residual_lfp.

Notes

The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.

Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.

The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1

The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))

Examples

>>> import numpy as np
>>> # Generate sample data and frequency arrays
>>> data = np.random.randn(10, 10000)
>>> fs = 1000.0  # 1 kHz sampling rate
>>> fscale = np.linspace(0, 90, 100)
>>> period = np.random.rand(10, 100)  # Mock PSD data
>>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]}
>>> features = get_psd_decay_features(data, fs, fscale, period, bands)
>>> print(features.columns)
Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error',
       'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta',
       'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')

References

Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.

ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]

Compute LF features from a numpy array.

Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,

’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].

Return type:

pd.DataFrame

Note

The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.

ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]

Compute CSD features from a numpy array.

Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.

Parameters:
  • data (np.ndarray) – Data array with shape (channels, samples).

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.

  • decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10. 1 skips scipy.signal.decimate entirely (passed through unchanged) rather than calling it with a no-op factor: scipy.signal.decimate(..., q=1, ...) designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, which scipy.signal.firwin rejects.

  • scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.

  • denoise (bool, optional) – Apply ibldsp.cadzow.cadzow_denoiser to the decimated data before the CSD finite difference. Defaults to True (today’s behavior). Set False when data has already been Cadzow-denoised upstream (e.g. combined with decimate=1 for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.

Returns:

DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,

’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].

Return type:

pd.DataFrame

Note

The function applies Cadzow denoising and current source density computation before computing the spectral features.

ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]

Compute AP features from a numpy array.

Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.

Parameters:
  • data (np.ndarray) – AP band data array with shape (channels, samples).

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • channel_labels (np.ndarray) – Array of channel quality labels.

Returns:

DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].

Return type:

pd.DataFrame

Raises:

AssertionError – If geometry or channel_labels are not provided.

Note

This function computes RMS values and cross-correlation ratios for action potential band data.

ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]

Perform spike detection using Dartsort.

This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.

Parameters:
  • data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.

  • fs (float) – Sampling frequency in Hz.

  • geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.

  • **params – Additional parameters for Dartsort configuration.

Returns:

A tuple containing:

  • df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.

  • d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.

Return type:

tuple

Note

This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.

DARTsort detects spikes on data normalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning: ptp and the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_val and the slopes computed by ibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.

ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]

Dartsort backend for spike detection.

This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Dartsort parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (DartParameters): Dartsort parameters object.

Return type:

tuple

ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]

SpikeInterface backend for spike detection.

This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – SpikeInterface parameters.

Returns:

A tuple containing:

  • df_spikes_ (pd.DataFrame): DataFrame with spike information.

  • d_waveforms (dict): Dictionary containing waveform data.

  • params_obj (dict): SpikeInterface parameters object.

Return type:

tuple

Raises:

Note

This function is currently a placeholder and needs full implementation for SpikeInterface backend support.

ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]

Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.

Parameters:
  • channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in d_waveforms["channel_index"]. Padded with the sentinel n_channels_real for unused neighbour slots.

  • spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g. df_spikes["channel"].

  • x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.

  • y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.

Returns:

(n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused

neighbour slots.

Return type:

np.ndarray

ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]

Spike detection and feature extraction with multiple backend support.

This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.

Parameters:
  • data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.

  • fs (int) – Sampling frequency in Hz.

  • geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.

  • return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.

  • backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.

  • scratch_dir (str, optional) – Directory for temporary files.

  • **params – Backend-specific parameters.

Returns:

If return_waveforms is False, returns DataFrame with

aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).

Return type:

pd.DataFrame or tuple

Raises:

ValueError – If an unknown backend is specified.

Note

The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.

ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]

Compute cross-correlation over auto-correlation ratio.

This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.

Parameters:
  • v (np.ndarray) – Voltage array for AP band with shape (nc, ns) where nc is number of channels and ns is number of samples.

  • geometry (dict) – Geometry dictionary with ‘x’ and ‘y’ arrays for electrode positions.

  • n_neighbor (int, optional) – Number of neighboring channels to consider. Defaults to 3.

Returns:

Array of size (nc,) containing the correlation ratios.

Return type:

np.ndarray

Note

The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.

ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]

Denoise AP features using total variation filter.

Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.

Parameters:
  • feature (np.ndarray) – AP feature to denoise with shape (nc,).

  • xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).

  • labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.

  • fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.

Returns:

Denoised AP features with shape (nc,).

Return type:

np.ndarray

Note

This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.

class ephysatlas.features._EphysTransformerInterface[source]

Bases: ABC, OneToOneFeatureMixin, TransformerMixin, BaseEstimator

Abstract base class for electrophysiological feature transformers.

This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.

Note

This is an abstract base class that should not be instantiated directly.

validate_X(X)[source]
Return type:

None

fit_transform(X=None, y=None)[source]

Fit the transformer and transform the data.

Parameters:
  • X (pd.DataFrame, optional) – Input data to fit and transform.

  • y – Ignored, present for compatibility with scikit-learn interface.

Returns:

Transformed data.

Return type:

pd.DataFrame

class ephysatlas.features.EphysTransformer[source]

Bases: _EphysTransformerInterface

fit(X=None, y=None)[source]
transform(X, y=None)[source]
class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Bases: _EphysTransformerInterface

__init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

TV-denoise electrophysiological features, with one weight factor per feature group.

Parameters

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.

channel_labelsnp.ndarray, optional

Channel quality annotation array with shape (nc,).

transform(X, y=None)[source]
fit(X=None, y=None)[source]
ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]

Applies total variation filter denoising to the features of a single probe insertion dataframe.

This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.

Parameters

df_pidpandas.DataFrame

DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.

facfloat or dict, default=DEFAULT_FAC

Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.

Returns

pandas.DataFrame

A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.