Features Module
The features module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings.
Module Overview
This module includes: * Feature extraction from AP (action potential) and LF (local field potential) bands * Current source density (CSD) computation * Spike detection and waveform analysis * Feature denoising using total variation filters * Data validation using Pandera schemas * Transformer classes for scikit-learn compatibility
Core Functions
Feature Computation
Electrophysiological feature extraction and processing module.
This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.
The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility
Classes
- DartParameters
Configuration parameters for Dartsort backend
- ChannelDataFrameSchema
Pandera schema for channel data validation
- ModelLfFeatures
Schema for local field potential features
- ModelCsdFeatures
Schema for current source density features
- ModelApFeatures
Schema for action potential features
- ModelSpikeFeatures
Schema for spike waveform features
- ModelSpikeShapeFeatures
Sparse, neuroscientist-facing remapping of ModelSpikeFeatures
- ModelChannelLayout
Schema for channel layout information
- ModelHistologyPlanned
Schema for planned histology coordinates
- ModelHistologyResolved
Schema for resolved histology coordinates
- ModelRawFeatures
Combined schema for all raw features
- EphysTransformer
Transformer for applying feature transformations
- EphysDenoiser
Transformer for denoising electrophysiological features
Functions
- _setup_scratch_directory
Set up scratch directory with fallback logic
- voltage_features_set
Get list of feature column names by provenance
- lf
Compute LF features from numpy array
- csd
Compute CSD features from numpy array
- ap
Compute AP features from numpy array
- dart_subtraction_numpy
Perform spike detection using Dartsort
- spikes
Spike detection and feature extraction with multiple backend support
- remap_waveform_shape_features
Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set
- remap_waveform_shape_features_volume
Same remapping, for a brainwide encoding volume instead of a dataframe
- xcor_acor_ratio
Compute cross-correlation over auto-correlation ratio
- denoise_shank
Denoise AP features using total variation filter
- denoise_dataframe
Apply total variation filter denoising to features
Constants
- __features_version__str
Version of the feature extractor code
- BANDSdict
Frequency bands for spectral analysis
- FEATURES_LISTlist
List of available feature types
Examples
>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000) # 64 channels, 30k samples
>>> fs = 30000 # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))
Notes
This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.
See Also
ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities
- ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]
Set up scratch directory with fallback logic.
- Parameters:
scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.
- Returns:
Path to the created scratch directory.
- Return type:
Path
Note
This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.
- ephysatlas.features.get_feature_cmin(feature_name)[source]
Get the minimum value for a given feature.
- Parameters:
feature_name (str) – Name of the feature.
- Returns:
Minimum value for the feature.
- Return type:
Note
This function is currently a placeholder and needs implementation.
- class ephysatlas.features.DartParameters(**data)[source]
Bases:
BaseModelConfiguration parameters for Dartsort backend.
This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.
- Variables:
localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.
chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).
trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.
scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.
n_jobs (int) – DARTsort worker count used to build its
ComputationConfig.0runs in the main process; positive values select that many workers. Defaults to 0.detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as
voltage_threshold. Defaults to 4.0.spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.
positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.
residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as
subtraction_threshold. Defaults to 3.162 (sqrt(10)).
- model_config: ClassVar[ConfigDict] = {}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]
Bases:
DataFrameModelPandera schema for channel data validation.
This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.
Note
x/y/zare in metres, whileaxial_um/lateral_umare in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.- Variables:
pid (Series[str]) – Probe insertion ID.
channel (Series[int]) – Channel index.
x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).
axial_um (Series[float]) – Axial distance in micrometers.
lateral_um (Series[float]) – Lateral distance in micrometers.
acronym (Series[str]) – Brain region acronym.
atlas_id (Series[int]) – Atlas region identifier.
- pid: str = 'pid'
- channel: int = 'channel'
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- acronym: str = 'acronym'
- atlas_id: int = 'atlas_id'
- class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]
Bases:
DataFrameModelBase class for channel-based feature schemas.
This is an abstract base class that provides the foundation for all channel-based feature validation schemas.
Note
The channel field is expected to be an index in derived classes.
- class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for local field potential features.
This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf (Series[float]) – Root mean square of LFP signal in dB.
psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).
psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).
psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).
psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).
psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).
psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).
- rms_lf: float = 'rms_lf'
- psd_lfp: float = 'psd_lfp'
- psd_delta: float = 'psd_delta'
- psd_theta: float = 'psd_theta'
- psd_alpha: float = 'psd_alpha'
- psd_beta: float = 'psd_beta'
- psd_gamma: float = 'psd_gamma'
- class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for current source density features.
This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.
psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).
psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).
psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).
psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).
psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).
psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).
- rms_lf_csd: float = 'rms_lf_csd'
- psd_delta_csd: float = 'psd_delta_csd'
- psd_theta_csd: float = 'psd_theta_csd'
- psd_alpha_csd: float = 'psd_alpha_csd'
- psd_beta_csd: float = 'psd_beta_csd'
- psd_gamma_csd: float = 'psd_gamma_csd'
- psd_lfp_csd: float = 'psd_lfp_csd'
- class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for action potential features.
This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.
- Variables:
- rms_ap: float = 'rms_ap'
- cor_ratio: float = 'cor_ratio'
- channel_labels: int = 'channel_labels'
- class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSpike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).
- Variables:
alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).
alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).
depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).
polarity (Series[float]) – Spike polarity, dimensionless.
recovery_slope (Series[float]) – Slope during recovery phase (V/s).
repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).
slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).
slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).
spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).
spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).
spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).
tip_val (Series[float]) – Tip amplitude value (V).
Note
tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.
- alpha_mean: float = 'alpha_mean'
- alpha_std: float = 'alpha_std'
- depolarisation_slope: float = 'depolarisation_slope'
- polarity: float = 'polarity'
- recovery_slope: float = 'recovery_slope'
- repolarisation_slope: float = 'repolarisation_slope'
- spike_count: float = 'spike_count'
- tip_val: float = 'tip_val'
- class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for spike waveform features.
This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.
- Variables:
peak_time_secs (Series[float]) – Time to peak in seconds.
peak_val (Series[float]) – Peak amplitude value (V).
recovery_time_secs (Series[float]) – Recovery time in seconds.
tip_time_secs (Series[float]) – Time to tip in seconds.
trough_time_secs (Series[float]) – Time to trough in seconds.
trough_val (Series[float]) – Trough amplitude value (V).
Note
peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.
- peak_time_secs: float = 'peak_time_secs'
- peak_val: float = 'peak_val'
- recovery_time_secs: float = 'recovery_time_secs'
- tip_time_secs: float = 'tip_time_secs'
- trough_time_secs: float = 'trough_time_secs'
- trough_val: float = 'trough_val'
- class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for the sparse, neuroscientist-facing spike-shape feature set.
Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.
The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.
- Variables:
- spike_width_secs: float = 'spike_width_secs'
- predepolarisation_width_secs: float = 'predepolarisation_width_secs'
- spike_amplitude: float = 'spike_amplitude'
- peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
- ephysatlas.features._remap_waveform_shape_arrays(get)[source]
Core of the waveform-shape remapping, agnostic to the container.
Parameters
- getcallable
get(name)returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).
Returns
- dict
Maps each ModelSpikeShapeFeatures column name to its array.
- ephysatlas.features.remap_waveform_shape_features(df)[source]
Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.
The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only
recovery_time_secs(exact constant offset oftrough_time_secs) andpeak_val/trough_val/peak_time_secs/trough_time_secs/tip_time_secs(reparametrised intospike_amplitude/peak_to_trough_ratio_log/spike_width_secs/predepolarisation_width_secswithout losing relative information) are actually dropped.tip_valand the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information.slowness_s_per_m/spatial_spread_um(#123) and their_stdcounterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.- Return type:
Parameters
- dfpandas.DataFrame
Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.
Returns
- pandas.DataFrame
Columns of ModelSpikeShapeFeatures, same index as df.
- ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]
Same transform as remap_waveform_shape_features, for a download_encoding_volume array.
Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.
Parameters
- ephys_atlas_volnumpy.ndarray
Shape
(nx, ny, nz, n_features).- feature_namesnumpy.ndarray or sequence of str
Length
n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.
Returns
- remapped_volnumpy.ndarray
Shape
(nx, ny, nz, 16), dtype float32.- remapped_feature_namesnumpy.ndarray
Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).
Raises
- KeyError
If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).
- class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for channel layout information.
This schema defines the structure and validation rules for channel layout features including spatial positioning.
- Variables:
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for planned histology coordinates.
This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.
Note
Same coordinate space as
ModelHistologyResolved(IBL, metres from Bregma); Populated byfeature_computation.add_target_coordinates, which queries Alyx trajectories withprovenance__lte,50and takes the best provenance available perPROVENANCE_PREFERENCE:Micro-manipulator(30), elsePlanned(10), elseHistology track(50); an insertion with none of the three raisesValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.- Variables:
- x_target: float = 'x_target'
- y_target: float = 'y_target'
- z_target: float = 'z_target'
- class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for resolved histology coordinates.
This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.
Note
Same coordinate space as
ModelHistologyPlanned. These come fromSpikeSortingLoader.load_channels()– the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the*_targetcolumns use, hence “at the most recent histology step”.- Variables:
x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).
atlas_id (Series[int]) – Atlas region identifier.
acronym (Series[str]) – Brain region acronym.
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- atlas_id: int = 'atlas_id'
- acronym: str = 'acronym'
- class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for all raw features.
This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.
Note
This class inherits from multiple feature schemas to provide a unified interface for all feature types.
- class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeShapeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for denoised features, after the waveform-shape remap.
Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.
Note
Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.
- class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]
Bases:
DataFrameModelSchema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.
- pid: str = 'pid'
- eid: str = 'eid'
- bwm: bool = 'bwm'
- ephysatlas.features.feature_group_of_column_map()[source]
Map each denoisable feature column name to its provenance group.
- Return type:
Returns
- dict
Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.
- ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]
Get list of feature column names by provenance.
This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.
- Parameters:
features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.
- Returns:
Sorted list of feature column names excluding the ‘channel’ column.
- Return type:
Note
The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.
- ephysatlas.features._get_power_in_band(fscale, period, band)[source]
Calculate power in a specific frequency band.
- Parameters:
fscale (np.ndarray) – Frequency scale array.
period (np.ndarray) – Periodogram values.
band (list) – Frequency band [low, high] in Hz.
- Returns:
Power in the specified band in dB relative to v/sqrt(Hz).
- Return type:
np.ndarray
Note
This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.
- ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]
Extract power spectral density decay features from electrophysiological data.
This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.
Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.
The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.
Parameters
- datanp.ndarray
Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.
- fsfloat
Sampling frequency of the data in Hz.
- fscalenp.ndarray
Frequency scale array corresponding to the period data.
- periodnp.ndarray
2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.
- bandsdict
Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.
- npersegint, optional
Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.
- PSD_rangelist of float, optional
Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.
Returns
- pd.DataFrame
One row per channel. Aperiodic-component columns:
aperiodic_offset(log10 power at 1 Hz),aperiodic_exponent(1/f slope in log-log space),decay_fit_error(RMS fit error),decay_fit_r_squared(goodness of fit) anddecay_n_peaks(number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal):psd_residual_delta,psd_residual_theta,psd_residual_alpha,psd_residual_beta,psd_residual_gammaandpsd_residual_lfp.
Notes
The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.
Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.
The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1
The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))
Examples
>>> import numpy as np >>> # Generate sample data and frequency arrays >>> data = np.random.randn(10, 10000) >>> fs = 1000.0 # 1 kHz sampling rate >>> fscale = np.linspace(0, 90, 100) >>> period = np.random.rand(10, 100) # Mock PSD data >>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]} >>> features = get_psd_decay_features(data, fs, fscale, period, bands) >>> print(features.columns) Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error', 'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta', 'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')
References
Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.
- ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]
Compute LF features from a numpy array.
Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,
’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].
- Return type:
pd.DataFrame
Note
The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.
- ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]
Compute CSD features from a numpy array.
Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
data (np.ndarray) – Data array with shape (channels, samples).
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.
decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10.
1skipsscipy.signal.decimateentirely (passed through unchanged) rather than calling it with a no-op factor:scipy.signal.decimate(..., q=1, ...)designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, whichscipy.signal.firwinrejects.scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.
denoise (bool, optional) – Apply
ibldsp.cadzow.cadzow_denoiserto the decimated data before the CSD finite difference. Defaults toTrue(today’s behavior). SetFalsewhendatahas already been Cadzow-denoised upstream (e.g. combined withdecimate=1for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,
’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].
- Return type:
pd.DataFrame
Note
The function applies Cadzow denoising and current source density computation before computing the spectral features.
- ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]
Compute AP features from a numpy array.
Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.
- Parameters:
data (np.ndarray) – AP band data array with shape (channels, samples).
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
channel_labels (np.ndarray) – Array of channel quality labels.
- Returns:
DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].
- Return type:
pd.DataFrame
- Raises:
AssertionError – If geometry or channel_labels are not provided.
Note
This function computes RMS values and cross-correlation ratios for action potential band data.
- ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]
Perform spike detection using Dartsort.
This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.
- Parameters:
data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
**params – Additional parameters for Dartsort configuration.
- Returns:
A tuple containing:
df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.
d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.
- Return type:
Note
This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.
DARTsort detects spikes on
datanormalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning:ptpand the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_valand the slopes computed byibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.
- ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]
Dartsort backend for spike detection.
This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (DartParameters): Dartsort parameters object.
- Return type:
- ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]
SpikeInterface backend for spike detection.
This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (dict): SpikeInterface parameters object.
- Return type:
- Raises:
ImportError – If SpikeInterface is not installed.
NotImplementedError – This function is not yet fully implemented.
Note
This function is currently a placeholder and needs full implementation for SpikeInterface backend support.
- ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]
Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.
- Parameters:
channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in
d_waveforms["channel_index"]. Padded with the sentineln_channels_realfor unused neighbour slots.spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g.
df_spikes["channel"].x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.
y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.
- Returns:
- (n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused
neighbour slots.
- Return type:
np.ndarray
- ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]
Spike detection and feature extraction with multiple backend support.
This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.
- Parameters:
data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.
fs (int) – Sampling frequency in Hz.
geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.
return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.
backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.
scratch_dir (str, optional) – Directory for temporary files.
**params – Backend-specific parameters.
- Returns:
- If return_waveforms is False, returns DataFrame with
aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).
- Return type:
pd.DataFrame or tuple
- Raises:
ValueError – If an unknown backend is specified.
Note
The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.
- ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]
Compute cross-correlation over auto-correlation ratio.
This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.
- Parameters:
- Returns:
Array of size (nc,) containing the correlation ratios.
- Return type:
np.ndarray
Note
The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.
- ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]
Denoise AP features using total variation filter.
Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.
- Parameters:
feature (np.ndarray) – AP feature to denoise with shape (nc,).
xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).
labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.
fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.
- Returns:
Denoised AP features with shape (nc,).
- Return type:
np.ndarray
Note
This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.
- class ephysatlas.features._EphysTransformerInterface[source]
Bases:
ABC,OneToOneFeatureMixin,TransformerMixin,BaseEstimatorAbstract base class for electrophysiological feature transformers.
This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.
Note
This is an abstract base class that should not be instantiated directly.
- fit_transform(X=None, y=None)[source]
Fit the transformer and transform the data.
- Parameters:
X (pd.DataFrame, optional) – Input data to fit and transform.
y – Ignored, present for compatibility with scikit-learn interface.
- Returns:
Transformed data.
- Return type:
pd.DataFrame
- class ephysatlas.features.EphysTransformer[source]
Bases:
_EphysTransformerInterface- fit(X=None, y=None)[source]
- transform(X, y=None)[source]
- class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Bases:
_EphysTransformerInterface- __init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
TV-denoise electrophysiological features, with one weight factor per feature group.
Parameters
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.
- channel_labelsnp.ndarray, optional
Channel quality annotation array with shape (nc,).
- transform(X, y=None)[source]
- fit(X=None, y=None)[source]
- ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Applies total variation filter denoising to the features of a single probe insertion dataframe.
This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.
Parameters
- df_pidpandas.DataFrame
DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.
Returns
- pandas.DataFrame
A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.
Data Classes
Electrophysiological feature extraction and processing module.
This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.
The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility
Classes
- DartParameters
Configuration parameters for Dartsort backend
- ChannelDataFrameSchema
Pandera schema for channel data validation
- ModelLfFeatures
Schema for local field potential features
- ModelCsdFeatures
Schema for current source density features
- ModelApFeatures
Schema for action potential features
- ModelSpikeFeatures
Schema for spike waveform features
- ModelSpikeShapeFeatures
Sparse, neuroscientist-facing remapping of ModelSpikeFeatures
- ModelChannelLayout
Schema for channel layout information
- ModelHistologyPlanned
Schema for planned histology coordinates
- ModelHistologyResolved
Schema for resolved histology coordinates
- ModelRawFeatures
Combined schema for all raw features
- EphysTransformer
Transformer for applying feature transformations
- EphysDenoiser
Transformer for denoising electrophysiological features
Functions
- _setup_scratch_directory
Set up scratch directory with fallback logic
- voltage_features_set
Get list of feature column names by provenance
- lf
Compute LF features from numpy array
- csd
Compute CSD features from numpy array
- ap
Compute AP features from numpy array
- dart_subtraction_numpy
Perform spike detection using Dartsort
- spikes
Spike detection and feature extraction with multiple backend support
- remap_waveform_shape_features
Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set
- remap_waveform_shape_features_volume
Same remapping, for a brainwide encoding volume instead of a dataframe
- xcor_acor_ratio
Compute cross-correlation over auto-correlation ratio
- denoise_shank
Denoise AP features using total variation filter
- denoise_dataframe
Apply total variation filter denoising to features
Constants
- __features_version__str
Version of the feature extractor code
- BANDSdict
Frequency bands for spectral analysis
- FEATURES_LISTlist
List of available feature types
Examples
>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000) # 64 channels, 30k samples
>>> fs = 30000 # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))
Notes
This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.
See Also
ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities
- ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]
Set up scratch directory with fallback logic.
- Parameters:
scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.
- Returns:
Path to the created scratch directory.
- Return type:
Path
Note
This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.
- ephysatlas.features.get_feature_cmin(feature_name)[source]
Get the minimum value for a given feature.
- Parameters:
feature_name (str) – Name of the feature.
- Returns:
Minimum value for the feature.
- Return type:
Note
This function is currently a placeholder and needs implementation.
- class ephysatlas.features.DartParameters(**data)[source]
Bases:
BaseModelConfiguration parameters for Dartsort backend.
This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.
- Variables:
localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.
chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).
trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.
scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.
n_jobs (int) – DARTsort worker count used to build its
ComputationConfig.0runs in the main process; positive values select that many workers. Defaults to 0.detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as
voltage_threshold. Defaults to 4.0.spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.
positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.
residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as
subtraction_threshold. Defaults to 3.162 (sqrt(10)).
- model_config: ClassVar[ConfigDict] = {}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]
Bases:
DataFrameModelPandera schema for channel data validation.
This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.
Note
x/y/zare in metres, whileaxial_um/lateral_umare in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.- Variables:
pid (Series[str]) – Probe insertion ID.
channel (Series[int]) – Channel index.
x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).
axial_um (Series[float]) – Axial distance in micrometers.
lateral_um (Series[float]) – Lateral distance in micrometers.
acronym (Series[str]) – Brain region acronym.
atlas_id (Series[int]) – Atlas region identifier.
- pid: str = 'pid'
- channel: int = 'channel'
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- acronym: str = 'acronym'
- atlas_id: int = 'atlas_id'
- class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]
Bases:
DataFrameModelBase class for channel-based feature schemas.
This is an abstract base class that provides the foundation for all channel-based feature validation schemas.
Note
The channel field is expected to be an index in derived classes.
- class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for local field potential features.
This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf (Series[float]) – Root mean square of LFP signal in dB.
psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).
psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).
psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).
psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).
psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).
psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).
- rms_lf: float = 'rms_lf'
- psd_lfp: float = 'psd_lfp'
- psd_delta: float = 'psd_delta'
- psd_theta: float = 'psd_theta'
- psd_alpha: float = 'psd_alpha'
- psd_beta: float = 'psd_beta'
- psd_gamma: float = 'psd_gamma'
- class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for current source density features.
This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.
psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).
psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).
psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).
psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).
psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).
psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).
- rms_lf_csd: float = 'rms_lf_csd'
- psd_delta_csd: float = 'psd_delta_csd'
- psd_theta_csd: float = 'psd_theta_csd'
- psd_alpha_csd: float = 'psd_alpha_csd'
- psd_beta_csd: float = 'psd_beta_csd'
- psd_gamma_csd: float = 'psd_gamma_csd'
- psd_lfp_csd: float = 'psd_lfp_csd'
- class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for action potential features.
This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.
- Variables:
- rms_ap: float = 'rms_ap'
- cor_ratio: float = 'cor_ratio'
- channel_labels: int = 'channel_labels'
- class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSpike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).
- Variables:
alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).
alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).
depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).
polarity (Series[float]) – Spike polarity, dimensionless.
recovery_slope (Series[float]) – Slope during recovery phase (V/s).
repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).
slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).
slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).
spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).
spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).
spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).
tip_val (Series[float]) – Tip amplitude value (V).
Note
tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.
- alpha_mean: float = 'alpha_mean'
- alpha_std: float = 'alpha_std'
- depolarisation_slope: float = 'depolarisation_slope'
- polarity: float = 'polarity'
- recovery_slope: float = 'recovery_slope'
- repolarisation_slope: float = 'repolarisation_slope'
- spike_count: float = 'spike_count'
- tip_val: float = 'tip_val'
- class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for spike waveform features.
This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.
- Variables:
peak_time_secs (Series[float]) – Time to peak in seconds.
peak_val (Series[float]) – Peak amplitude value (V).
recovery_time_secs (Series[float]) – Recovery time in seconds.
tip_time_secs (Series[float]) – Time to tip in seconds.
trough_time_secs (Series[float]) – Time to trough in seconds.
trough_val (Series[float]) – Trough amplitude value (V).
Note
peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.
- peak_time_secs: float = 'peak_time_secs'
- peak_val: float = 'peak_val'
- recovery_time_secs: float = 'recovery_time_secs'
- tip_time_secs: float = 'tip_time_secs'
- trough_time_secs: float = 'trough_time_secs'
- trough_val: float = 'trough_val'
- class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for the sparse, neuroscientist-facing spike-shape feature set.
Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.
The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.
- Variables:
- spike_width_secs: float = 'spike_width_secs'
- predepolarisation_width_secs: float = 'predepolarisation_width_secs'
- spike_amplitude: float = 'spike_amplitude'
- peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
- ephysatlas.features._remap_waveform_shape_arrays(get)[source]
Core of the waveform-shape remapping, agnostic to the container.
Parameters
- getcallable
get(name)returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).
Returns
- dict
Maps each ModelSpikeShapeFeatures column name to its array.
- ephysatlas.features.remap_waveform_shape_features(df)[source]
Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.
The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only
recovery_time_secs(exact constant offset oftrough_time_secs) andpeak_val/trough_val/peak_time_secs/trough_time_secs/tip_time_secs(reparametrised intospike_amplitude/peak_to_trough_ratio_log/spike_width_secs/predepolarisation_width_secswithout losing relative information) are actually dropped.tip_valand the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information.slowness_s_per_m/spatial_spread_um(#123) and their_stdcounterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.- Return type:
Parameters
- dfpandas.DataFrame
Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.
Returns
- pandas.DataFrame
Columns of ModelSpikeShapeFeatures, same index as df.
- ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]
Same transform as remap_waveform_shape_features, for a download_encoding_volume array.
Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.
Parameters
- ephys_atlas_volnumpy.ndarray
Shape
(nx, ny, nz, n_features).- feature_namesnumpy.ndarray or sequence of str
Length
n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.
Returns
- remapped_volnumpy.ndarray
Shape
(nx, ny, nz, 16), dtype float32.- remapped_feature_namesnumpy.ndarray
Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).
Raises
- KeyError
If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).
- class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for channel layout information.
This schema defines the structure and validation rules for channel layout features including spatial positioning.
- Variables:
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for planned histology coordinates.
This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.
Note
Same coordinate space as
ModelHistologyResolved(IBL, metres from Bregma); Populated byfeature_computation.add_target_coordinates, which queries Alyx trajectories withprovenance__lte,50and takes the best provenance available perPROVENANCE_PREFERENCE:Micro-manipulator(30), elsePlanned(10), elseHistology track(50); an insertion with none of the three raisesValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.- Variables:
- x_target: float = 'x_target'
- y_target: float = 'y_target'
- z_target: float = 'z_target'
- class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for resolved histology coordinates.
This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.
Note
Same coordinate space as
ModelHistologyPlanned. These come fromSpikeSortingLoader.load_channels()– the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the*_targetcolumns use, hence “at the most recent histology step”.- Variables:
x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).
atlas_id (Series[int]) – Atlas region identifier.
acronym (Series[str]) – Brain region acronym.
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- atlas_id: int = 'atlas_id'
- acronym: str = 'acronym'
- class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for all raw features.
This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.
Note
This class inherits from multiple feature schemas to provide a unified interface for all feature types.
- class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeShapeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for denoised features, after the waveform-shape remap.
Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.
Note
Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.
- class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]
Bases:
DataFrameModelSchema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.
- pid: str = 'pid'
- eid: str = 'eid'
- bwm: bool = 'bwm'
- ephysatlas.features.feature_group_of_column_map()[source]
Map each denoisable feature column name to its provenance group.
- Return type:
Returns
- dict
Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.
- ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]
Get list of feature column names by provenance.
This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.
- Parameters:
features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.
- Returns:
Sorted list of feature column names excluding the ‘channel’ column.
- Return type:
Note
The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.
- ephysatlas.features._get_power_in_band(fscale, period, band)[source]
Calculate power in a specific frequency band.
- Parameters:
fscale (np.ndarray) – Frequency scale array.
period (np.ndarray) – Periodogram values.
band (list) – Frequency band [low, high] in Hz.
- Returns:
Power in the specified band in dB relative to v/sqrt(Hz).
- Return type:
np.ndarray
Note
This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.
- ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]
Extract power spectral density decay features from electrophysiological data.
This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.
Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.
The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.
Parameters
- datanp.ndarray
Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.
- fsfloat
Sampling frequency of the data in Hz.
- fscalenp.ndarray
Frequency scale array corresponding to the period data.
- periodnp.ndarray
2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.
- bandsdict
Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.
- npersegint, optional
Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.
- PSD_rangelist of float, optional
Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.
Returns
- pd.DataFrame
One row per channel. Aperiodic-component columns:
aperiodic_offset(log10 power at 1 Hz),aperiodic_exponent(1/f slope in log-log space),decay_fit_error(RMS fit error),decay_fit_r_squared(goodness of fit) anddecay_n_peaks(number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal):psd_residual_delta,psd_residual_theta,psd_residual_alpha,psd_residual_beta,psd_residual_gammaandpsd_residual_lfp.
Notes
The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.
Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.
The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1
The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))
Examples
>>> import numpy as np >>> # Generate sample data and frequency arrays >>> data = np.random.randn(10, 10000) >>> fs = 1000.0 # 1 kHz sampling rate >>> fscale = np.linspace(0, 90, 100) >>> period = np.random.rand(10, 100) # Mock PSD data >>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]} >>> features = get_psd_decay_features(data, fs, fscale, period, bands) >>> print(features.columns) Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error', 'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta', 'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')
References
Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.
- ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]
Compute LF features from a numpy array.
Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,
’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].
- Return type:
pd.DataFrame
Note
The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.
- ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]
Compute CSD features from a numpy array.
Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
data (np.ndarray) – Data array with shape (channels, samples).
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.
decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10.
1skipsscipy.signal.decimateentirely (passed through unchanged) rather than calling it with a no-op factor:scipy.signal.decimate(..., q=1, ...)designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, whichscipy.signal.firwinrejects.scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.
denoise (bool, optional) – Apply
ibldsp.cadzow.cadzow_denoiserto the decimated data before the CSD finite difference. Defaults toTrue(today’s behavior). SetFalsewhendatahas already been Cadzow-denoised upstream (e.g. combined withdecimate=1for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,
’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].
- Return type:
pd.DataFrame
Note
The function applies Cadzow denoising and current source density computation before computing the spectral features.
- ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]
Compute AP features from a numpy array.
Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.
- Parameters:
data (np.ndarray) – AP band data array with shape (channels, samples).
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
channel_labels (np.ndarray) – Array of channel quality labels.
- Returns:
DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].
- Return type:
pd.DataFrame
- Raises:
AssertionError – If geometry or channel_labels are not provided.
Note
This function computes RMS values and cross-correlation ratios for action potential band data.
- ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]
Perform spike detection using Dartsort.
This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.
- Parameters:
data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
**params – Additional parameters for Dartsort configuration.
- Returns:
A tuple containing:
df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.
d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.
- Return type:
Note
This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.
DARTsort detects spikes on
datanormalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning:ptpand the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_valand the slopes computed byibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.
- ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]
Dartsort backend for spike detection.
This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (DartParameters): Dartsort parameters object.
- Return type:
- ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]
SpikeInterface backend for spike detection.
This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (dict): SpikeInterface parameters object.
- Return type:
- Raises:
ImportError – If SpikeInterface is not installed.
NotImplementedError – This function is not yet fully implemented.
Note
This function is currently a placeholder and needs full implementation for SpikeInterface backend support.
- ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]
Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.
- Parameters:
channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in
d_waveforms["channel_index"]. Padded with the sentineln_channels_realfor unused neighbour slots.spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g.
df_spikes["channel"].x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.
y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.
- Returns:
- (n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused
neighbour slots.
- Return type:
np.ndarray
- ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]
Spike detection and feature extraction with multiple backend support.
This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.
- Parameters:
data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.
fs (int) – Sampling frequency in Hz.
geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.
return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.
backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.
scratch_dir (str, optional) – Directory for temporary files.
**params – Backend-specific parameters.
- Returns:
- If return_waveforms is False, returns DataFrame with
aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).
- Return type:
pd.DataFrame or tuple
- Raises:
ValueError – If an unknown backend is specified.
Note
The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.
- ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]
Compute cross-correlation over auto-correlation ratio.
This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.
- Parameters:
- Returns:
Array of size (nc,) containing the correlation ratios.
- Return type:
np.ndarray
Note
The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.
- ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]
Denoise AP features using total variation filter.
Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.
- Parameters:
feature (np.ndarray) – AP feature to denoise with shape (nc,).
xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).
labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.
fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.
- Returns:
Denoised AP features with shape (nc,).
- Return type:
np.ndarray
Note
This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.
- class ephysatlas.features._EphysTransformerInterface[source]
Bases:
ABC,OneToOneFeatureMixin,TransformerMixin,BaseEstimatorAbstract base class for electrophysiological feature transformers.
This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.
Note
This is an abstract base class that should not be instantiated directly.
- fit_transform(X=None, y=None)[source]
Fit the transformer and transform the data.
- Parameters:
X (pd.DataFrame, optional) – Input data to fit and transform.
y – Ignored, present for compatibility with scikit-learn interface.
- Returns:
Transformed data.
- Return type:
pd.DataFrame
- class ephysatlas.features.EphysTransformer[source]
Bases:
_EphysTransformerInterface- fit(X=None, y=None)[source]
- transform(X, y=None)[source]
- class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Bases:
_EphysTransformerInterface- __init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
TV-denoise electrophysiological features, with one weight factor per feature group.
Parameters
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.
- channel_labelsnp.ndarray, optional
Channel quality annotation array with shape (nc,).
- transform(X, y=None)[source]
- fit(X=None, y=None)[source]
- ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Applies total variation filter denoising to the features of a single probe insertion dataframe.
This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.
Parameters
- df_pidpandas.DataFrame
DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.
Returns
- pandas.DataFrame
A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.
Transformers
Electrophysiological feature extraction and processing module.
This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.
The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility
Classes
- DartParameters
Configuration parameters for Dartsort backend
- ChannelDataFrameSchema
Pandera schema for channel data validation
- ModelLfFeatures
Schema for local field potential features
- ModelCsdFeatures
Schema for current source density features
- ModelApFeatures
Schema for action potential features
- ModelSpikeFeatures
Schema for spike waveform features
- ModelSpikeShapeFeatures
Sparse, neuroscientist-facing remapping of ModelSpikeFeatures
- ModelChannelLayout
Schema for channel layout information
- ModelHistologyPlanned
Schema for planned histology coordinates
- ModelHistologyResolved
Schema for resolved histology coordinates
- ModelRawFeatures
Combined schema for all raw features
- EphysTransformer
Transformer for applying feature transformations
- EphysDenoiser
Transformer for denoising electrophysiological features
Functions
- _setup_scratch_directory
Set up scratch directory with fallback logic
- voltage_features_set
Get list of feature column names by provenance
- lf
Compute LF features from numpy array
- csd
Compute CSD features from numpy array
- ap
Compute AP features from numpy array
- dart_subtraction_numpy
Perform spike detection using Dartsort
- spikes
Spike detection and feature extraction with multiple backend support
- remap_waveform_shape_features
Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set
- remap_waveform_shape_features_volume
Same remapping, for a brainwide encoding volume instead of a dataframe
- xcor_acor_ratio
Compute cross-correlation over auto-correlation ratio
- denoise_shank
Denoise AP features using total variation filter
- denoise_dataframe
Apply total variation filter denoising to features
Constants
- __features_version__str
Version of the feature extractor code
- BANDSdict
Frequency bands for spectral analysis
- FEATURES_LISTlist
List of available feature types
Examples
>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000) # 64 channels, 30k samples
>>> fs = 30000 # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))
Notes
This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.
See Also
ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities
- ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]
Set up scratch directory with fallback logic.
- Parameters:
scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.
- Returns:
Path to the created scratch directory.
- Return type:
Path
Note
This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.
- ephysatlas.features.get_feature_cmin(feature_name)[source]
Get the minimum value for a given feature.
- Parameters:
feature_name (str) – Name of the feature.
- Returns:
Minimum value for the feature.
- Return type:
Note
This function is currently a placeholder and needs implementation.
- class ephysatlas.features.DartParameters(**data)[source]
Bases:
BaseModelConfiguration parameters for Dartsort backend.
This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.
- Variables:
localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.
chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).
trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.
scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.
n_jobs (int) – DARTsort worker count used to build its
ComputationConfig.0runs in the main process; positive values select that many workers. Defaults to 0.detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as
voltage_threshold. Defaults to 4.0.spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.
positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.
residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as
subtraction_threshold. Defaults to 3.162 (sqrt(10)).
- model_config: ClassVar[ConfigDict] = {}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]
Bases:
DataFrameModelPandera schema for channel data validation.
This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.
Note
x/y/zare in metres, whileaxial_um/lateral_umare in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.- Variables:
pid (Series[str]) – Probe insertion ID.
channel (Series[int]) – Channel index.
x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).
axial_um (Series[float]) – Axial distance in micrometers.
lateral_um (Series[float]) – Lateral distance in micrometers.
acronym (Series[str]) – Brain region acronym.
atlas_id (Series[int]) – Atlas region identifier.
- pid: str = 'pid'
- channel: int = 'channel'
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- acronym: str = 'acronym'
- atlas_id: int = 'atlas_id'
- class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]
Bases:
DataFrameModelBase class for channel-based feature schemas.
This is an abstract base class that provides the foundation for all channel-based feature validation schemas.
Note
The channel field is expected to be an index in derived classes.
- class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for local field potential features.
This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf (Series[float]) – Root mean square of LFP signal in dB.
psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).
psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).
psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).
psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).
psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).
psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).
- rms_lf: float = 'rms_lf'
- psd_lfp: float = 'psd_lfp'
- psd_delta: float = 'psd_delta'
- psd_theta: float = 'psd_theta'
- psd_alpha: float = 'psd_alpha'
- psd_beta: float = 'psd_beta'
- psd_gamma: float = 'psd_gamma'
- class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for current source density features.
This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.
psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).
psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).
psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).
psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).
psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).
psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).
- rms_lf_csd: float = 'rms_lf_csd'
- psd_delta_csd: float = 'psd_delta_csd'
- psd_theta_csd: float = 'psd_theta_csd'
- psd_alpha_csd: float = 'psd_alpha_csd'
- psd_beta_csd: float = 'psd_beta_csd'
- psd_gamma_csd: float = 'psd_gamma_csd'
- psd_lfp_csd: float = 'psd_lfp_csd'
- class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for action potential features.
This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.
- Variables:
- rms_ap: float = 'rms_ap'
- cor_ratio: float = 'cor_ratio'
- channel_labels: int = 'channel_labels'
- class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSpike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).
- Variables:
alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).
alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).
depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).
polarity (Series[float]) – Spike polarity, dimensionless.
recovery_slope (Series[float]) – Slope during recovery phase (V/s).
repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).
slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).
slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).
spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).
spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).
spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).
tip_val (Series[float]) – Tip amplitude value (V).
Note
tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.
- alpha_mean: float = 'alpha_mean'
- alpha_std: float = 'alpha_std'
- depolarisation_slope: float = 'depolarisation_slope'
- polarity: float = 'polarity'
- recovery_slope: float = 'recovery_slope'
- repolarisation_slope: float = 'repolarisation_slope'
- spike_count: float = 'spike_count'
- tip_val: float = 'tip_val'
- class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for spike waveform features.
This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.
- Variables:
peak_time_secs (Series[float]) – Time to peak in seconds.
peak_val (Series[float]) – Peak amplitude value (V).
recovery_time_secs (Series[float]) – Recovery time in seconds.
tip_time_secs (Series[float]) – Time to tip in seconds.
trough_time_secs (Series[float]) – Time to trough in seconds.
trough_val (Series[float]) – Trough amplitude value (V).
Note
peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.
- peak_time_secs: float = 'peak_time_secs'
- peak_val: float = 'peak_val'
- recovery_time_secs: float = 'recovery_time_secs'
- tip_time_secs: float = 'tip_time_secs'
- trough_time_secs: float = 'trough_time_secs'
- trough_val: float = 'trough_val'
- class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for the sparse, neuroscientist-facing spike-shape feature set.
Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.
The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.
- Variables:
- spike_width_secs: float = 'spike_width_secs'
- predepolarisation_width_secs: float = 'predepolarisation_width_secs'
- spike_amplitude: float = 'spike_amplitude'
- peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
- ephysatlas.features._remap_waveform_shape_arrays(get)[source]
Core of the waveform-shape remapping, agnostic to the container.
Parameters
- getcallable
get(name)returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).
Returns
- dict
Maps each ModelSpikeShapeFeatures column name to its array.
- ephysatlas.features.remap_waveform_shape_features(df)[source]
Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.
The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only
recovery_time_secs(exact constant offset oftrough_time_secs) andpeak_val/trough_val/peak_time_secs/trough_time_secs/tip_time_secs(reparametrised intospike_amplitude/peak_to_trough_ratio_log/spike_width_secs/predepolarisation_width_secswithout losing relative information) are actually dropped.tip_valand the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information.slowness_s_per_m/spatial_spread_um(#123) and their_stdcounterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.- Return type:
Parameters
- dfpandas.DataFrame
Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.
Returns
- pandas.DataFrame
Columns of ModelSpikeShapeFeatures, same index as df.
- ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]
Same transform as remap_waveform_shape_features, for a download_encoding_volume array.
Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.
Parameters
- ephys_atlas_volnumpy.ndarray
Shape
(nx, ny, nz, n_features).- feature_namesnumpy.ndarray or sequence of str
Length
n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.
Returns
- remapped_volnumpy.ndarray
Shape
(nx, ny, nz, 16), dtype float32.- remapped_feature_namesnumpy.ndarray
Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).
Raises
- KeyError
If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).
- class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for channel layout information.
This schema defines the structure and validation rules for channel layout features including spatial positioning.
- Variables:
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for planned histology coordinates.
This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.
Note
Same coordinate space as
ModelHistologyResolved(IBL, metres from Bregma); Populated byfeature_computation.add_target_coordinates, which queries Alyx trajectories withprovenance__lte,50and takes the best provenance available perPROVENANCE_PREFERENCE:Micro-manipulator(30), elsePlanned(10), elseHistology track(50); an insertion with none of the three raisesValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.- Variables:
- x_target: float = 'x_target'
- y_target: float = 'y_target'
- z_target: float = 'z_target'
- class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for resolved histology coordinates.
This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.
Note
Same coordinate space as
ModelHistologyPlanned. These come fromSpikeSortingLoader.load_channels()– the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the*_targetcolumns use, hence “at the most recent histology step”.- Variables:
x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).
atlas_id (Series[int]) – Atlas region identifier.
acronym (Series[str]) – Brain region acronym.
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- atlas_id: int = 'atlas_id'
- acronym: str = 'acronym'
- class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for all raw features.
This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.
Note
This class inherits from multiple feature schemas to provide a unified interface for all feature types.
- class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeShapeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for denoised features, after the waveform-shape remap.
Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.
Note
Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.
- class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]
Bases:
DataFrameModelSchema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.
- pid: str = 'pid'
- eid: str = 'eid'
- bwm: bool = 'bwm'
- ephysatlas.features.feature_group_of_column_map()[source]
Map each denoisable feature column name to its provenance group.
- Return type:
Returns
- dict
Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.
- ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]
Get list of feature column names by provenance.
This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.
- Parameters:
features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.
- Returns:
Sorted list of feature column names excluding the ‘channel’ column.
- Return type:
Note
The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.
- ephysatlas.features._get_power_in_band(fscale, period, band)[source]
Calculate power in a specific frequency band.
- Parameters:
fscale (np.ndarray) – Frequency scale array.
period (np.ndarray) – Periodogram values.
band (list) – Frequency band [low, high] in Hz.
- Returns:
Power in the specified band in dB relative to v/sqrt(Hz).
- Return type:
np.ndarray
Note
This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.
- ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]
Extract power spectral density decay features from electrophysiological data.
This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.
Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.
The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.
Parameters
- datanp.ndarray
Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.
- fsfloat
Sampling frequency of the data in Hz.
- fscalenp.ndarray
Frequency scale array corresponding to the period data.
- periodnp.ndarray
2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.
- bandsdict
Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.
- npersegint, optional
Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.
- PSD_rangelist of float, optional
Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.
Returns
- pd.DataFrame
One row per channel. Aperiodic-component columns:
aperiodic_offset(log10 power at 1 Hz),aperiodic_exponent(1/f slope in log-log space),decay_fit_error(RMS fit error),decay_fit_r_squared(goodness of fit) anddecay_n_peaks(number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal):psd_residual_delta,psd_residual_theta,psd_residual_alpha,psd_residual_beta,psd_residual_gammaandpsd_residual_lfp.
Notes
The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.
Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.
The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1
The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))
Examples
>>> import numpy as np >>> # Generate sample data and frequency arrays >>> data = np.random.randn(10, 10000) >>> fs = 1000.0 # 1 kHz sampling rate >>> fscale = np.linspace(0, 90, 100) >>> period = np.random.rand(10, 100) # Mock PSD data >>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]} >>> features = get_psd_decay_features(data, fs, fscale, period, bands) >>> print(features.columns) Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error', 'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta', 'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')
References
Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.
- ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]
Compute LF features from a numpy array.
Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,
’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].
- Return type:
pd.DataFrame
Note
The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.
- ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]
Compute CSD features from a numpy array.
Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
data (np.ndarray) – Data array with shape (channels, samples).
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.
decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10.
1skipsscipy.signal.decimateentirely (passed through unchanged) rather than calling it with a no-op factor:scipy.signal.decimate(..., q=1, ...)designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, whichscipy.signal.firwinrejects.scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.
denoise (bool, optional) – Apply
ibldsp.cadzow.cadzow_denoiserto the decimated data before the CSD finite difference. Defaults toTrue(today’s behavior). SetFalsewhendatahas already been Cadzow-denoised upstream (e.g. combined withdecimate=1for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,
’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].
- Return type:
pd.DataFrame
Note
The function applies Cadzow denoising and current source density computation before computing the spectral features.
- ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]
Compute AP features from a numpy array.
Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.
- Parameters:
data (np.ndarray) – AP band data array with shape (channels, samples).
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
channel_labels (np.ndarray) – Array of channel quality labels.
- Returns:
DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].
- Return type:
pd.DataFrame
- Raises:
AssertionError – If geometry or channel_labels are not provided.
Note
This function computes RMS values and cross-correlation ratios for action potential band data.
- ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]
Perform spike detection using Dartsort.
This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.
- Parameters:
data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
**params – Additional parameters for Dartsort configuration.
- Returns:
A tuple containing:
df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.
d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.
- Return type:
Note
This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.
DARTsort detects spikes on
datanormalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning:ptpand the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_valand the slopes computed byibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.
- ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]
Dartsort backend for spike detection.
This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (DartParameters): Dartsort parameters object.
- Return type:
- ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]
SpikeInterface backend for spike detection.
This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (dict): SpikeInterface parameters object.
- Return type:
- Raises:
ImportError – If SpikeInterface is not installed.
NotImplementedError – This function is not yet fully implemented.
Note
This function is currently a placeholder and needs full implementation for SpikeInterface backend support.
- ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]
Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.
- Parameters:
channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in
d_waveforms["channel_index"]. Padded with the sentineln_channels_realfor unused neighbour slots.spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g.
df_spikes["channel"].x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.
y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.
- Returns:
- (n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused
neighbour slots.
- Return type:
np.ndarray
- ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]
Spike detection and feature extraction with multiple backend support.
This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.
- Parameters:
data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.
fs (int) – Sampling frequency in Hz.
geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.
return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.
backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.
scratch_dir (str, optional) – Directory for temporary files.
**params – Backend-specific parameters.
- Returns:
- If return_waveforms is False, returns DataFrame with
aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).
- Return type:
pd.DataFrame or tuple
- Raises:
ValueError – If an unknown backend is specified.
Note
The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.
- ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]
Compute cross-correlation over auto-correlation ratio.
This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.
- Parameters:
- Returns:
Array of size (nc,) containing the correlation ratios.
- Return type:
np.ndarray
Note
The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.
- ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]
Denoise AP features using total variation filter.
Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.
- Parameters:
feature (np.ndarray) – AP feature to denoise with shape (nc,).
xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).
labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.
fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.
- Returns:
Denoised AP features with shape (nc,).
- Return type:
np.ndarray
Note
This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.
- class ephysatlas.features._EphysTransformerInterface[source]
Bases:
ABC,OneToOneFeatureMixin,TransformerMixin,BaseEstimatorAbstract base class for electrophysiological feature transformers.
This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.
Note
This is an abstract base class that should not be instantiated directly.
- fit_transform(X=None, y=None)[source]
Fit the transformer and transform the data.
- Parameters:
X (pd.DataFrame, optional) – Input data to fit and transform.
y – Ignored, present for compatibility with scikit-learn interface.
- Returns:
Transformed data.
- Return type:
pd.DataFrame
- class ephysatlas.features.EphysTransformer[source]
Bases:
_EphysTransformerInterface- fit(X=None, y=None)[source]
- transform(X, y=None)[source]
- class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Bases:
_EphysTransformerInterface- __init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
TV-denoise electrophysiological features, with one weight factor per feature group.
Parameters
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.
- channel_labelsnp.ndarray, optional
Channel quality annotation array with shape (nc,).
- transform(X, y=None)[source]
- fit(X=None, y=None)[source]
- ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Applies total variation filter denoising to the features of a single probe insertion dataframe.
This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.
Parameters
- df_pidpandas.DataFrame
DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.
Returns
- pandas.DataFrame
A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.
Configuration
Electrophysiological feature extraction and processing module.
This module provides comprehensive tools for extracting, processing, and analyzing electrophysiological features from neural recordings. It supports multiple backends for spike detection and feature computation, including Dartsort and SpikeInterface.
The module includes: - Feature extraction from AP (action potential) and LF (local field potential) bands - Current source density (CSD) computation - Spike detection and waveform analysis - Feature denoising using total variation filters - Data validation using Pandera schemas - Transformer classes for scikit-learn compatibility
Classes
- DartParameters
Configuration parameters for Dartsort backend
- ChannelDataFrameSchema
Pandera schema for channel data validation
- ModelLfFeatures
Schema for local field potential features
- ModelCsdFeatures
Schema for current source density features
- ModelApFeatures
Schema for action potential features
- ModelSpikeFeatures
Schema for spike waveform features
- ModelSpikeShapeFeatures
Sparse, neuroscientist-facing remapping of ModelSpikeFeatures
- ModelChannelLayout
Schema for channel layout information
- ModelHistologyPlanned
Schema for planned histology coordinates
- ModelHistologyResolved
Schema for resolved histology coordinates
- ModelRawFeatures
Combined schema for all raw features
- EphysTransformer
Transformer for applying feature transformations
- EphysDenoiser
Transformer for denoising electrophysiological features
Functions
- _setup_scratch_directory
Set up scratch directory with fallback logic
- voltage_features_set
Get list of feature column names by provenance
- lf
Compute LF features from numpy array
- csd
Compute CSD features from numpy array
- ap
Compute AP features from numpy array
- dart_subtraction_numpy
Perform spike detection using Dartsort
- spikes
Spike detection and feature extraction with multiple backend support
- remap_waveform_shape_features
Remap ModelSpikeFeatures’ 18 raw waveform columns onto a sparser set
- remap_waveform_shape_features_volume
Same remapping, for a brainwide encoding volume instead of a dataframe
- xcor_acor_ratio
Compute cross-correlation over auto-correlation ratio
- denoise_shank
Denoise AP features using total variation filter
- denoise_dataframe
Apply total variation filter denoising to features
Constants
- __features_version__str
Version of the feature extractor code
- BANDSdict
Frequency bands for spectral analysis
- FEATURES_LISTlist
List of available feature types
Examples
>>> from ephysatlas.features import lf, csd, ap
>>> import numpy as np
>>>
>>> # Generate sample data
>>> data = np.random.randn(64, 30000) # 64 channels, 30k samples
>>> fs = 30000 # 30 kHz sampling rate
>>>
>>> # Compute LF features
>>> lf_features = lf(data, fs)
>>>
>>> # Compute CSD features
>>> geometry = {'x': np.arange(64), 'y': np.zeros(64)}
>>> csd_features = csd(data, fs, geometry)
>>>
>>> # Compute AP features
>>> ap_data = np.random.randn(64, 10000)
>>> ap_features = ap(ap_data, geometry, np.zeros(64))
Notes
This module requires several dependencies including numpy, pandas, scipy, scikit-image, and optionally dartsort for advanced spike detection. GPU acceleration is supported through the dartsort backend.
See Also
ibldsp.waveforms : Waveform processing utilities ibldsp.cadzow : Cadzow denoising algorithms ibldsp.voltage : Voltage processing utilities
- ephysatlas.features._setup_scratch_directory(scratch_dir=None)[source]
Set up scratch directory with fallback logic.
- Parameters:
scratch_dir (Path or str, optional) – Preferred scratch directory path. If None, will try system defaults.
- Returns:
Path to the created scratch directory.
- Return type:
Path
Note
This function first tries to use the SDSC scratch directory (/scratch/dartsort/), and falls back to the system temp directory if that fails.
- ephysatlas.features.get_feature_cmin(feature_name)[source]
Get the minimum value for a given feature.
- Parameters:
feature_name (str) – Name of the feature.
- Returns:
Minimum value for the feature.
- Return type:
Note
This function is currently a placeholder and needs implementation.
- class ephysatlas.features.DartParameters(**data)[source]
Bases:
BaseModelConfiguration parameters for Dartsort backend.
This class defines the parameters used for spike detection and feature extraction using the Dartsort algorithm.
- Variables:
localization_radius (float) – Radius in micrometers for spike localization. Defaults to 150.
chunk_length_samples (int) – Length of data chunks in samples for processing. Defaults to 2^15 (32768).
trough_offset (int) – Offset in samples from spike peak to trough. Defaults to 42.
scratch_dir (Path or str, optional) – Scratch directory for temporary files. If None, will use system defaults.
n_jobs (int) – DARTsort worker count used to build its
ComputationConfig.0runs in the main process; positive values select that many workers. Defaults to 0.detection_threshold (float) – Peak detection threshold, in units of channel RMS. Passed to DARTsort as
voltage_threshold. Defaults to 4.0.spatial_dedup_radius_um (float) – Radius in micrometres within which only the largest simultaneous peak is kept. Defaults to 150.0.
positive_temporal_dedup_radius_samples (int) – Samples around a trough in which positive peaks are suppressed. Defaults to 7.
residnorm_decrease_threshold (float) – Minimum residual-norm decrease for a detected spike to be subtracted. Passed to DARTsort as
subtraction_threshold. Defaults to 3.162 (sqrt(10)).
- model_config: ClassVar[ConfigDict] = {}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- class ephysatlas.features.ChannelDataFrameSchema(*args, **kwargs)[source]
Bases:
DataFrameModelPandera schema for channel data validation.
This schema defines the structure and validation rules for channel information including spatial coordinates and anatomical labels.
Note
x/y/zare in metres, whileaxial_um/lateral_umare in micrometres. The two length units coexist in one table: the coordinates follow the IBL/iblatlas convention (metres from Bregma), whereas the probe geometry is expressed in the micrometres.- Variables:
pid (Series[str]) – Probe insertion ID.
channel (Series[int]) – Channel index.
x (Series[float]) – X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Z-coordinate in metres from Bregma (IBL coordinates space).
axial_um (Series[float]) – Axial distance in micrometers.
lateral_um (Series[float]) – Lateral distance in micrometers.
acronym (Series[str]) – Brain region acronym.
atlas_id (Series[int]) – Atlas region identifier.
- pid: str = 'pid'
- channel: int = 'channel'
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- acronym: str = 'acronym'
- atlas_id: int = 'atlas_id'
- class ephysatlas.features.BaseChannelFeatures(*args, **kwargs)[source]
Bases:
DataFrameModelBase class for channel-based feature schemas.
This is an abstract base class that provides the foundation for all channel-based feature validation schemas.
Note
The channel field is expected to be an index in derived classes.
- class ephysatlas.features.ModelLfFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for local field potential features.
This schema defines the structure and validation rules for local field potential (LFP) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf (Series[float]) – Root mean square of LFP signal in dB.
psd_delta (Series[float]) – Power spectral density in delta band (0-4 Hz).
psd_theta (Series[float]) – Power spectral density in theta band (4-10 Hz).
psd_alpha (Series[float]) – Power spectral density in alpha band (8-12 Hz).
psd_beta (Series[float]) – Power spectral density in beta band (15-30 Hz).
psd_gamma (Series[float]) – Power spectral density in gamma band (30-90 Hz).
psd_lfp (Series[float]) – Power spectral density in full LFP band (0-90 Hz).
- rms_lf: float = 'rms_lf'
- psd_lfp: float = 'psd_lfp'
- psd_delta: float = 'psd_delta'
- psd_theta: float = 'psd_theta'
- psd_alpha: float = 'psd_alpha'
- psd_beta: float = 'psd_beta'
- psd_gamma: float = 'psd_gamma'
- class ephysatlas.features.ModelCsdFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for current source density features.
This schema defines the structure and validation rules for current source density (CSD) features including RMS values and power spectral density across different frequency bands.
- Variables:
rms_lf_csd (Series[float]) – Root mean square of CSD signal in dB.
psd_delta_csd (Series[float]) – CSD power spectral density in delta band (0-4 Hz).
psd_theta_csd (Series[float]) – CSD power spectral density in theta band (4-10 Hz).
psd_alpha_csd (Series[float]) – CSD power spectral density in alpha band (8-12 Hz).
psd_beta_csd (Series[float]) – CSD power spectral density in beta band (15-30 Hz).
psd_gamma_csd (Series[float]) – CSD power spectral density in gamma band (30-90 Hz).
psd_lfp_csd (Series[float]) – CSD power spectral density in full LFP band (0-90 Hz).
- rms_lf_csd: float = 'rms_lf_csd'
- psd_delta_csd: float = 'psd_delta_csd'
- psd_theta_csd: float = 'psd_theta_csd'
- psd_alpha_csd: float = 'psd_alpha_csd'
- psd_beta_csd: float = 'psd_beta_csd'
- psd_gamma_csd: float = 'psd_gamma_csd'
- psd_lfp_csd: float = 'psd_lfp_csd'
- class ephysatlas.features.ModelApFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for action potential features.
This schema defines the structure and validation rules for action potential (AP) features including RMS values and correlation ratios.
- Variables:
- rms_ap: float = 'rms_ap'
- cor_ratio: float = 'cor_ratio'
- channel_labels: int = 'channel_labels'
- class ephysatlas.features.ModelSpikeSharedFeatures(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSpike-shape fields shared verbatim by ModelSpikeFeatures and ModelSpikeShapeFeatures (the sparse remap keeps these unchanged).
- Variables:
alpha_mean (Series[float]) – Mean alpha parameter for spike localization (N/A).
alpha_std (Series[float]) – Standard deviation of alpha parameter (N/A).
depolarisation_slope (Series[float]) – Slope during depolarization phase (V/s).
polarity (Series[float]) – Spike polarity, dimensionless.
recovery_slope (Series[float]) – Slope during recovery phase (V/s).
repolarisation_slope (Series[float]) – Slope during repolarization phase (V/s).
slowness_s_per_m (Series[float]) – Signed axial slowness of the multi-channel waveform (s/m).
slowness_s_per_m_std (Series[float]) – Across-spike standard deviation of slowness_s_per_m on the channel (s/m).
spatial_spread_um (Series[float]) – Amplitude-weighted spatial spread of the multi-channel waveform (um).
spatial_spread_um_std (Series[float]) – Across-spike standard deviation of spatial_spread_um on the channel (um).
spike_count (Series[float]) – Spike rate, spikes/s (log2 transformed).
tip_val (Series[float]) – Tip amplitude value (V).
Note
tip_val and the 3 slope columns are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts (DARTsort z-scores internally by channel RMS for detection only). alpha_mean/alpha_std are DARTsort’s localisation amplitude, fit jointly across a multichannel snippet whose channels each carry a different RMS; there is no single scale factor to recover Volts from them, so they stay in DARTsort’s native, non-physical units. slowness_s_per_m/spatial_spread_um are also weighted by these Volts-scale amplitudes (see ibldsp.waveforms.compute_slowness/ compute_spatial_spread), but the values themselves – and those of their _std counterparts – are geometric (seconds per metre / micrometres), not amplitudes.
- alpha_mean: float = 'alpha_mean'
- alpha_std: float = 'alpha_std'
- depolarisation_slope: float = 'depolarisation_slope'
- polarity: float = 'polarity'
- recovery_slope: float = 'recovery_slope'
- repolarisation_slope: float = 'repolarisation_slope'
- spike_count: float = 'spike_count'
- tip_val: float = 'tip_val'
- class ephysatlas.features.ModelSpikeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for spike waveform features.
This schema defines the structure and validation rules for spike waveform features including timing, amplitude, and slope characteristics. Adds the raw per-spike-event columns to ModelSpikeSharedFeatures.
- Variables:
peak_time_secs (Series[float]) – Time to peak in seconds.
peak_val (Series[float]) – Peak amplitude value (V).
recovery_time_secs (Series[float]) – Recovery time in seconds.
tip_time_secs (Series[float]) – Time to tip in seconds.
trough_time_secs (Series[float]) – Time to trough in seconds.
trough_val (Series[float]) – Trough amplitude value (V).
Note
peak_val/trough_val are computed from waveforms that dart_subtraction_numpy has already rescaled back to Volts; see ModelSpikeSharedFeatures for tip_val/the slopes/alpha.
- peak_time_secs: float = 'peak_time_secs'
- peak_val: float = 'peak_val'
- recovery_time_secs: float = 'recovery_time_secs'
- tip_time_secs: float = 'tip_time_secs'
- trough_time_secs: float = 'trough_time_secs'
- trough_val: float = 'trough_val'
- class ephysatlas.features.ModelSpikeShapeFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeSharedFeaturesSchema for the sparse, neuroscientist-facing spike-shape feature set.
Output of remap_waveform_shape_features. Adds the reparametrised columns to ModelSpikeSharedFeatures’s unchanged pass-through fields. Note this schema’s ‘peak’ is the negative deflection and ‘trough’ the positive rebound after it, opposite common neurophysiology usage, so spike_width_secs is the classic trough-to-peak width despite the name order.
The 3 inherited slope columns correlate strongly with the amplitude/duration columns below (r=0.96-0.99 with algebraic reconstructions from them) and add no real extra information, but are kept precomputed anyway since they’re a commonly-used, handy quantity on their own.
- Variables:
- spike_width_secs: float = 'spike_width_secs'
- predepolarisation_width_secs: float = 'predepolarisation_width_secs'
- spike_amplitude: float = 'spike_amplitude'
- peak_to_trough_ratio_log: float = 'peak_to_trough_ratio_log'
- ephysatlas.features._remap_waveform_shape_arrays(get)[source]
Core of the waveform-shape remapping, agnostic to the container.
Parameters
- getcallable
get(name)returns the raw ModelSpikeFeatures column/slice for name (a pandas.Series or an (nx, ny, nz) array both work: only elementwise +, -, /, np.log, np.abs are used below).
Returns
- dict
Maps each ModelSpikeShapeFeatures column name to its array.
- ephysatlas.features.remap_waveform_shape_features(df)[source]
Remap ModelSpikeFeatures’s 18 raw waveform columns onto the sparser ModelSpikeShapeFeatures set.
The 14 columns the PCA was run over are redundant: it needs only ~7-8 components for 95% of their variance. Only
recovery_time_secs(exact constant offset oftrough_time_secs) andpeak_val/trough_val/peak_time_secs/trough_time_secs/tip_time_secs(reparametrised intospike_amplitude/peak_to_trough_ratio_log/spike_width_secs/predepolarisation_width_secswithout losing relative information) are actually dropped.tip_valand the 3 slopes (depolarisation_slope/repolarisation_slope/recovery_slope) are kept unchanged despite correlating with the amplitude/duration columns (r=0.79-0.99) – handy precomputed quantities in their own right, even if not strictly independent information.slowness_s_per_m/spatial_spread_um(#123) and their_stdcounterparts (#127) are geometric, postdate that PCA redundancy analysis and were not part of it, and are also kept unchanged.- Return type:
Parameters
- dfpandas.DataFrame
Must contain the columns of ModelSpikeFeatures (e.g. the output of spike, raw or denoised). slowness_s_per_m/spatial_spread_um (#123) and their _std counterparts (#127), all nullable, are treated as all-NaN if altogether absent, so data saved before they existed can still be remapped.
Returns
- pandas.DataFrame
Columns of ModelSpikeShapeFeatures, same index as df.
- ephysatlas.features.remap_waveform_shape_features_volume(ephys_atlas_vol, feature_names)[source]
Same transform as remap_waveform_shape_features, for a download_encoding_volume array.
Values there are unnormalised already, so no de-z-scoring is needed. Does not derive mean_per_feature/std_per_feature for the new columns – those come from the original per-channel dataset, not the volume itself.
Parameters
- ephys_atlas_volnumpy.ndarray
Shape
(nx, ny, nz, n_features).- feature_namesnumpy.ndarray or sequence of str
Length
n_features, naming ephys_atlas_vol’s last axis; must include the 18 raw columns of ModelSpikeFeatures.
Returns
- remapped_volnumpy.ndarray
Shape
(nx, ny, nz, 16), dtype float32.- remapped_feature_namesnumpy.ndarray
Length 16, naming remapped_vol’s last axis (ModelSpikeShapeFeatures column order).
Raises
- KeyError
If a required raw feature name is absent from feature_names (except slowness_s_per_m/spatial_spread_um, #123, and their _std counterparts, #127, all nullable: treated as all-NaN if altogether absent, so volumes computed before they existed can still be remapped).
- class ephysatlas.features.ModelChannelLayout(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for channel layout information.
This schema defines the structure and validation rules for channel layout features including spatial positioning.
- Variables:
- axial_um: float = 'axial_um'
- lateral_um: float = 'lateral_um'
- class ephysatlas.features.ModelHistologyPlanned(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for planned histology coordinates.
This schema defines the structure and validation rules for planned histology coordinates before actual histological analysis. The coordinates come from the best trajectory Alyx holds for the insertion – micro-manipulator, else planned, else histology track – see feature_computation.PROVENANCE_PREFERENCE.
Note
Same coordinate space as
ModelHistologyResolved(IBL, metres from Bregma); Populated byfeature_computation.add_target_coordinates, which queries Alyx trajectories withprovenance__lte,50and takes the best provenance available perPROVENANCE_PREFERENCE:Micro-manipulator(30), elsePlanned(10), elseHistology track(50); an insertion with none of the three raisesValueError. A micro-manipulator or planned trajectory is converted from the needles (in-vivo) atlas into Allen space before being interpolated along the track; a histology track is already in Allen space and is interpolated as is.- Variables:
- x_target: float = 'x_target'
- y_target: float = 'y_target'
- z_target: float = 'z_target'
- class ephysatlas.features.ModelHistologyResolved(*args, **kwargs)[source]
Bases:
BaseChannelFeaturesSchema for resolved histology coordinates.
This schema defines the structure and validation rules for resolved histology coordinates after actual histological analysis.
Note
Same coordinate space as
ModelHistologyPlanned. These come fromSpikeSortingLoader.load_channels()– the highest-provenance alignment Alyx holds for the insertion, above the micro-manipulator estimate the*_targetcolumns use, hence “at the most recent histology step”.- Variables:
x (Series[float]) – Resolved X-coordinate in metres from Bregma (IBL coordinates space).
y (Series[float]) – Resolved Y-coordinate in metres from Bregma (IBL coordinates space).
z (Series[float]) – Resolved Z-coordinate in metres from Bregma (IBL coordinates space).
atlas_id (Series[int]) – Atlas region identifier.
acronym (Series[str]) – Brain region acronym.
- x: float = 'x'
- y: float = 'y'
- z: float = 'z'
- atlas_id: int = 'atlas_id'
- acronym: str = 'acronym'
- class ephysatlas.features.ModelRawFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for all raw features.
This schema combines all individual feature schemas into a single comprehensive schema for raw electrophysiological data validation.
Note
This class inherits from multiple feature schemas to provide a unified interface for all feature types.
- class ephysatlas.features.ModelDenoisedFeatures(*args, **kwargs)[source]
Bases:
ModelSpikeShapeFeatures,ModelCsdFeatures,ModelApFeatures,ModelLfFeatures,ModelChannelLayoutCombined schema for denoised features, after the waveform-shape remap.
Identical to ModelRawFeatures apart from the spike block: denoise_raw_features_data applies remap_waveform_shape_features, which drops ModelSpikeFeatures’ 6 raw timing/amplitude columns in favour of ModelSpikeShapeFeatures’ 4 reparametrised ones.
Note
Denoised tables written before that remap still match ModelRawFeatures, so ephysatlas.data.read_features_from_disk chooses between the two schemas on the columns actually present rather than on which file it loaded.
- class ephysatlas.features.ModelProbeDetails(*args, **kwargs)[source]
Bases:
DataFrameModelSchema for probe insertion metadata (df_probe_details.pqt). One row per probe insertion.
- pid: str = 'pid'
- eid: str = 'eid'
- bwm: bool = 'bwm'
- ephysatlas.features.feature_group_of_column_map()[source]
Map each denoisable feature column name to its provenance group.
- Return type:
Returns
- dict
Maps feature column name to one of ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’.
- ephysatlas.features.voltage_features_set(features_list=['raw_ap', 'raw_lf', 'localisation', 'waveforms'])[source]
Get list of feature column names by provenance.
This function returns the list of features columns names depending on their provenance. This is useful to select the columns for training.
- Parameters:
features_list (list, optional) – List of feature groups to include. Defaults to [‘raw_ap’, ‘raw_lf’, ‘localisation’, ‘waveforms’]. Use ‘all’ to include all available feature groups.
- Returns:
Sorted list of feature column names excluding the ‘channel’ column.
- Return type:
Note
The looping preserves the order of the features groups in the list. Available feature groups: ‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’, ‘micro-manipulator’.
- ephysatlas.features._get_power_in_band(fscale, period, band)[source]
Calculate power in a specific frequency band.
- Parameters:
fscale (np.ndarray) – Frequency scale array.
period (np.ndarray) – Periodogram values.
band (list) – Frequency band [low, high] in Hz.
- Returns:
Power in the specified band in dB relative to v/sqrt(Hz).
- Return type:
np.ndarray
Note
This function weights the frequencies using a cosine window and computes the weighted average power in the specified band.
- ephysatlas.features.get_psd_decay_features(data, fs, fscale, period, bands, nperseg=2048, PSD_range=[0, 90])[source]
Extract power spectral density decay features from electrophysiological data.
This function computes spectral parameterization features that characterize the aperiodic (1/f) component of the power spectral density (PSD) using the specparam library. It fits a model to separate periodic peaks from the aperiodic background in the frequency domain, providing insights into the underlying neural dynamics.
Additionally, it computes residual power features by removing the fitted aperiodic component from the observed PSD, highlighting periodic components across different frequency bands.
The aperiodic component of neural signals is thought to reflect the balance of excitation and inhibition in neural circuits, making these features particularly useful for characterizing brain states and pathological conditions.
Parameters
- datanp.ndarray
Input electrophysiological data with shape (n_channels, n_samples). Each row represents a different recording channel.
- fsfloat
Sampling frequency of the data in Hz.
- fscalenp.ndarray
Frequency scale array corresponding to the period data.
- periodnp.ndarray
2D array of power spectral density values with shape (n_channels, n_frequencies). Each row represents the PSD for a different recording channel.
- bandsdict
Dictionary mapping frequency band names to [min_freq, max_freq] ranges in Hz. Used for computing residual power in specific frequency bands.
- npersegint, optional
Length of each segment for Welch’s method PSD estimation, by default 2048. Larger values provide better frequency resolution but less temporal averaging.
- PSD_rangelist of float, optional
Frequency range [min_freq, max_freq] in Hz for spectral parameterization, by default BANDS[“lfp”] which is [0, 90] Hz.
Returns
- pd.DataFrame
One row per channel. Aperiodic-component columns:
aperiodic_offset(log10 power at 1 Hz),aperiodic_exponent(1/f slope in log-log space),decay_fit_error(RMS fit error),decay_fit_r_squared(goodness of fit) anddecay_n_peaks(number of periodic peaks). Residual-power columns (periodic power per band after aperiodic removal):psd_residual_delta,psd_residual_theta,psd_residual_alpha,psd_residual_beta,psd_residual_gammaandpsd_residual_lfp.
Notes
The function uses the specparam library (formerly FOOOF) to separate periodic and aperiodic components of the PSD. The aperiodic component follows a 1/f^β relationship where β is the aperiodic exponent.
Channels with R-squared < 0.9 are flagged as having poor fits, which may indicate artifacts or unusual spectral properties.
The spectral model is configured with: - Peak width limits: [10, 15] Hz - Maximum peaks: 4 - Minimum peak height: 0.1
The residual curve is computed as: residual = 10^(log10(observed_PSD) - log10(fitted_aperiodic_component))
Examples
>>> import numpy as np >>> # Generate sample data and frequency arrays >>> data = np.random.randn(10, 10000) >>> fs = 1000.0 # 1 kHz sampling rate >>> fscale = np.linspace(0, 90, 100) >>> period = np.random.rand(10, 100) # Mock PSD data >>> bands = {'delta': [0, 4], 'theta': [4, 10], 'alpha': [8, 12]} >>> features = get_psd_decay_features(data, fs, fscale, period, bands) >>> print(features.columns) Index(['aperiodic_offset', 'aperiodic_exponent', 'decay_fit_error', 'decay_fit_r_squared', 'decay_n_peaks', 'psd_residual_delta', 'psd_residual_theta', 'psd_residual_alpha', ...], dtype='object')
References
Donoghue, T., Haller, M., Peterson, E. J., Varma, P., Sebastian, P., Gao, R., et al. (2020). Parameterizing neural power spectra into periodic and aperiodic components. Nature Neuroscience, 23(12), 1655-1665.
- ephysatlas.features.lf(data, fs, bands=None, decay_features=True)[source]
Compute LF features from a numpy array.
Computes the local field potential (LF) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf’, ‘psd_delta’,
’psd_theta’, ‘psd_alpha’, ‘psd_beta’, ‘psd_gamma’, ‘psd_lfp’].
- Return type:
pd.DataFrame
Note
The function computes RMS values and power spectral density for each frequency band defined in the BANDS constant.
- ephysatlas.features.csd(data, fs, geometry, bands=None, decimate=10, scale=True, denoise=True)[source]
Compute CSD features from a numpy array.
Computes the current source density (CSD) features from electrophysiological data including RMS values and power spectral density across different frequency bands.
- Parameters:
data (np.ndarray) – Data array with shape (channels, samples).
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
bands (dict, optional) – Dictionary with frequency bands to compute. Defaults to BANDS constant.
decimate (int, optional) – Decimation factor for CSD calculation. Defaults to 10.
1skipsscipy.signal.decimateentirely (passed through unchanged) rather than calling it with a no-op factor:scipy.signal.decimate(..., q=1, ...)designs an anti-aliasing filter with cutoff exactly at the Nyquist boundary, whichscipy.signal.firwinrejects.scale (bool, optional) – Forwarded to current_source_density. If True, scale the finite difference by the intercontact distance and tissue conductivity; if False, return the raw numerical finite difference. Defaults to True.
denoise (bool, optional) – Apply
ibldsp.cadzow.cadzow_denoiserto the decimated data before the CSD finite difference. Defaults toTrue(today’s behavior). SetFalsewhendatahas already been Cadzow-denoised upstream (e.g. combined withdecimate=1for a source that is already at the target rate), so it isn’t denoised a second time on top of a lossy reconstruction.
- Returns:
- DataFrame with columns [‘channel’, ‘rms_lf_csd’, ‘psd_delta_csd’,
’psd_theta_csd’, ‘psd_alpha_csd’, ‘psd_beta_csd’, ‘psd_gamma_csd’, ‘psd_lfp_csd’].
- Return type:
pd.DataFrame
Note
The function applies Cadzow denoising and current source density computation before computing the spectral features.
- ephysatlas.features.ap(data, geometry=None, channel_labels=None)[source]
Compute AP features from a numpy array.
Computes the action potential (AP) features from electrophysiological data including RMS values and correlation ratios.
- Parameters:
data (np.ndarray) – AP band data array with shape (channels, samples).
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
channel_labels (np.ndarray) – Array of channel quality labels.
- Returns:
DataFrame with columns [‘channel’, ‘rms_ap’, ‘cor_ratio’, ‘channel_labels’].
- Return type:
pd.DataFrame
- Raises:
AssertionError – If geometry or channel_labels are not provided.
Note
This function computes RMS values and cross-correlation ratios for action potential band data.
- ephysatlas.features.dart_subtraction_numpy(data, fs, geometry, params=None, scratch_dir=None, **extra)[source]
Perform spike detection using Dartsort.
This function performs spike detection and feature extraction using the Dartsort algorithm with configurable parameters.
- Parameters:
data (np.ndarray) – Voltage traces array with shape [nc, ns] in Volts, where nc is number of channels and ns is number of samples.
fs (float) – Sampling frequency in Hz.
geometry (dict) – Dictionary with channel geometry containing ‘x’ and ‘y’ arrays.
**params – Additional parameters for Dartsort configuration.
- Returns:
A tuple containing:
df_spikes (pd.DataFrame): DataFrame with spike information including sample indices, channels, peak-to-peak amplitudes (Volts), and localizations.
d_waveforms (dict): Dictionary containing raw and denoised waveforms (Volts) and channel indices.
- Return type:
Note
This function requires the dartsort package to be installed. It creates temporary directories for processing and cleans them up afterward. GPU acceleration is supported when available.
DARTsort detects spikes on
datanormalised by its own per-channel RMS (a fixed detection threshold in units of channel RMS), but that scaling is un-applied before returning:ptpand the waveform snippets are rescaled back to Volts using the same per-channel RMS, so every amplitude derived from them downstream (peak_val/trough_val/tip_valand the slopes computed byibldsp.waveforms.compute_spike_features) comes out in real units. The localisation amplitude (alpha) is the one exception: it is fit jointly across a multichannel snippet whose channels each carry a different RMS, so there is no single scale factor to un-apply and it is left in DARTsort’s native (arbitrary) units.
- ephysatlas.features._spikes_dartsort(data, fs, geometry, scratch_dir=None, **params)[source]
Dartsort backend for spike detection.
This function serves as the Dartsort backend for the main spikes function, handling spike detection and feature extraction using Dartsort.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (DartParameters): Dartsort parameters object.
- Return type:
- ephysatlas.features._spikes_spikeinterface(data, fs, geometry, scratch_dir=None, **params)[source]
SpikeInterface backend for spike detection.
This function serves as the SpikeInterface backend for the main spikes function, handling spike detection and feature extraction using SpikeInterface.
- Parameters:
- Returns:
A tuple containing:
df_spikes_(pd.DataFrame): DataFrame with spike information.d_waveforms (dict): Dictionary containing waveform data.
params_obj (dict): SpikeInterface parameters object.
- Return type:
- Raises:
ImportError – If SpikeInterface is not installed.
NotImplementedError – This function is not yet fully implemented.
Note
This function is currently a placeholder and needs full implementation for SpikeInterface backend support.
- ephysatlas.features._neighbour_channel_xy_um(channel_index, spike_channels, x_um, y_um)[source]
Per-spike (x, y) coordinates (um) of a waveform array’s neighbour channels.
- Parameters:
channel_index (np.ndarray) – (n_channels_real, n_neighbours) lookup from a real detection channel to its neighbour real-channel ids, as returned by DARTsort in
d_waveforms["channel_index"]. Padded with the sentineln_channels_realfor unused neighbour slots.spike_channels (np.ndarray) – (n_spikes,) detection channel (real channel id) of each spike, e.g.
df_spikes["channel"].x_um (np.ndarray) – (n_channels_real,) real channel x-coordinate, um.
y_um (np.ndarray) – (n_channels_real,) real channel y-coordinate, um.
- Returns:
- (n_spikes, n_neighbours, 2) coordinates, NaN for padded/unused
neighbour slots.
- Return type:
np.ndarray
- ephysatlas.features.spikes(data, fs, geometry, return_waveforms=True, backend='dartsort', scratch_dir=None, **params)[source]
Spike detection and feature extraction with multiple backend support.
This function performs spike detection and feature extraction using either Dartsort or SpikeInterface backend, with comprehensive feature computation including waveform analysis and spike characterization.
- Parameters:
data (np.ndarray) – Raw electrophysiology data with shape [nc, ns] where nc is number of channels and ns is number of samples.
fs (int) – Sampling frequency in Hz.
geometry (dict) – Channel geometry dictionary with ‘x’ and ‘y’ arrays.
return_waveforms (bool, optional) – Whether to return waveforms dictionary. Defaults to True.
backend (str, optional) – Backend to use (‘dartsort’ or ‘spikeinterface’). Defaults to ‘dartsort’.
scratch_dir (str, optional) – Directory for temporary files.
**params – Backend-specific parameters.
- Returns:
- If return_waveforms is False, returns DataFrame with
aggregated spike features per channel. If True, returns tuple of (DataFrame, waveforms_dict).
- Return type:
pd.DataFrame or tuple
- Raises:
ValueError – If an unknown backend is specified.
Note
The function aggregates spike features by channel and computes various waveform characteristics including timing, amplitude, and slope features. Both backends produce compatible output formats for further processing.
- ephysatlas.features.xcor_acor_ratio(v, geometry, n_neighbor=3)[source]
Compute cross-correlation over auto-correlation ratio.
This function calculates the ratio of cross-correlation between neighboring channels over the auto-correlation for each channel in the AP band data.
- Parameters:
- Returns:
Array of size (nc,) containing the correlation ratios.
- Return type:
np.ndarray
Note
The function computes covariance matrices and extracts diagonal elements to calculate cross-correlation ratios for neighboring channels.
- ephysatlas.features.denoise_shank(feature, xy, labels=None, fac=1)[source]
Denoise AP features using total variation filter.
Denoise the AP feature using a maximum variation filter. Interpolates the feature in a square grid, performs the filtering, and then interpolates back to the original grid.
- Parameters:
feature (np.ndarray) – AP feature to denoise with shape (nc,).
xy (np.ndarray) – Coordinates of the AP feature with shape (nc, 2).
labels (np.ndarray, optional) – Channel quality annotation array with shape (nc,). If different than 0, channel is discarded and interpolated. Set to None for no annotation. Defaults to None.
fac (int, optional) – Factor for the TV denoising in median deviation units. Defaults to 1.
- Returns:
Denoised AP features with shape (nc,).
- Return type:
np.ndarray
Note
This function uses scikit-image’s total variation Chambolle denoising algorithm to smooth the feature values while preserving edges.
- class ephysatlas.features._EphysTransformerInterface[source]
Bases:
ABC,OneToOneFeatureMixin,TransformerMixin,BaseEstimatorAbstract base class for electrophysiological feature transformers.
This class provides the interface for transformers that work with electrophysiological features, implementing scikit-learn’s transformer interface and setting pandas as the default output format.
Note
This is an abstract base class that should not be instantiated directly.
- fit_transform(X=None, y=None)[source]
Fit the transformer and transform the data.
- Parameters:
X (pd.DataFrame, optional) – Input data to fit and transform.
y – Ignored, present for compatibility with scikit-learn interface.
- Returns:
Transformed data.
- Return type:
pd.DataFrame
- class ephysatlas.features.EphysTransformer[source]
Bases:
_EphysTransformerInterface- fit(X=None, y=None)[source]
- transform(X, y=None)[source]
- class ephysatlas.features.EphysDenoiser(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Bases:
_EphysTransformerInterface- __init__(fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
TV-denoise electrophysiological features, with one weight factor per feature group.
Parameters
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor. Groups not specified in the dict default to 1.
- channel_labelsnp.ndarray, optional
Channel quality annotation array with shape (nc,).
- transform(X, y=None)[source]
- fit(X=None, y=None)[source]
- ephysatlas.features.denoise_dataframe(df_pid, fac={'raw_ap': 0.1, 'raw_lf': 0.1, 'raw_lf_csd': 0.1, 'waveforms': 3}, channel_labels=None)[source]
Applies total variation filter denoising to the features of a single probe insertion dataframe.
This function processes electrophysiological features by applying a total variation filter to denoise them. If a transformation is defined in the metadata schema for a feature, it will be applied before denoising. Channels marked with non-zero labels are treated as invalid and their values are interpolated from neighboring channels.
Parameters
- df_pidpandas.DataFrame
DataFrame containing probe insertion data with features to denoise. Must contain ‘lateral_um’, ‘axial_um’, and ‘labels’ columns.
- facfloat or dict, default=DEFAULT_FAC
Factor for the TV denoising in median deviation units. Higher values result in stronger denoising. Either a single scalar applied to every feature group, or a dict mapping a subset of {‘raw_ap’, ‘raw_lf’, ‘raw_lf_csd’, ‘waveforms’} to their own factor (see feature_group_of_column_map). Groups absent from the dict default to 1.
Returns
- pandas.DataFrame
A new dataframe with the same structure as the input, but with denoised feature values. Non-feature columns are copied without modification.