Utils Module
The utils module provides utility functions for data management, file operations, and metadata handling.
Module Overview
This module includes: * Directory and output management utilities * Metadata handling for parquet files * Data aggregation and processing functions * File operation helpers
Core Functions
Directory Management
Utility functions for electrophysiological data processing and file management.
This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.
The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation
Functions
- setup_output_directory
Set up hierarchical output directory structure for probe and snippet data
- get_aggregated_snippets_df
Aggregate metadata from all snippets in a probe-level directory
- add_metadata_to_parquet_files
Add metadata attributes to all Parquet files in a snippet directory
- _update_parquet_metadata
Update metadata for a single Parquet file (internal helper function)
Examples
>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
... 'output_dir': '/data/output',
... 'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... 't_start': 300.0,
... 'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
... base_level_dir='/data/output',
... snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
... pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... t_start=300.0,
... duration=5.0
... )
Notes
This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.
See Also
pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations
- ephysatlas.utils.setup_output_directory(params)[source]
Set up the output directory structure for probe and snippet data.
This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.
- Parameters:
params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming
- Returns:
A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory
If output_dir is None, returns (None, None).
- Return type:
tuple[Path, Path]
- Raises:
ValueError – If neither pid nor filename is provided.
Note
The function creates a hierarchical structure:
Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration
Example output structure:
|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c | |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0 | `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0 `-- af0a0534-9cdc-4a29-93c0-1342891d74ec |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0 `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
- ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]
Get a dataframe of metadata info for all snippets in the probe level directory.
This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.
- Parameters:
probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.
- Returns:
- DataFrame containing metadata from all snippets, with each
row representing one snippet and columns representing different metadata attributes.
- Return type:
pd.DataFrame
Note
The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.
- ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]
Add metadata attributes to all Parquet files in a snippet-level directory.
This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.
- Parameters:
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.
- Returns:
The function modifies files in place and does not return any values.
- Return type:
None
Note
The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.
Example
>>> add_metadata_to_parquet_files( ... base_level_dir='/data/probe1', ... snippet_level_dir='snippet_001', ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
- ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]
Update metadata attributes for a single Parquet file.
This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.
- Parameters:
file_path (Path) – Path to the Parquet file to be updated.
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.
- Returns:
The function modifies the file in place and does not return any values.
- Return type:
None
Note
The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.
Example
>>> _update_parquet_metadata( ... Path('data.pqt'), ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
Metadata Management
Utility functions for electrophysiological data processing and file management.
This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.
The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation
Functions
- setup_output_directory
Set up hierarchical output directory structure for probe and snippet data
- get_aggregated_snippets_df
Aggregate metadata from all snippets in a probe-level directory
- add_metadata_to_parquet_files
Add metadata attributes to all Parquet files in a snippet directory
- _update_parquet_metadata
Update metadata for a single Parquet file (internal helper function)
Examples
>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
... 'output_dir': '/data/output',
... 'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... 't_start': 300.0,
... 'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
... base_level_dir='/data/output',
... snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
... pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... t_start=300.0,
... duration=5.0
... )
Notes
This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.
See Also
pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations
- ephysatlas.utils.setup_output_directory(params)[source]
Set up the output directory structure for probe and snippet data.
This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.
- Parameters:
params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming
- Returns:
A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory
If output_dir is None, returns (None, None).
- Return type:
tuple[Path, Path]
- Raises:
ValueError – If neither pid nor filename is provided.
Note
The function creates a hierarchical structure:
Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration
Example output structure:
|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c | |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0 | `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0 `-- af0a0534-9cdc-4a29-93c0-1342891d74ec |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0 `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
- ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]
Get a dataframe of metadata info for all snippets in the probe level directory.
This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.
- Parameters:
probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.
- Returns:
- DataFrame containing metadata from all snippets, with each
row representing one snippet and columns representing different metadata attributes.
- Return type:
pd.DataFrame
Note
The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.
- ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]
Add metadata attributes to all Parquet files in a snippet-level directory.
This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.
- Parameters:
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.
- Returns:
The function modifies files in place and does not return any values.
- Return type:
None
Note
The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.
Example
>>> add_metadata_to_parquet_files( ... base_level_dir='/data/probe1', ... snippet_level_dir='snippet_001', ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
- ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]
Update metadata attributes for a single Parquet file.
This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.
- Parameters:
file_path (Path) – Path to the Parquet file to be updated.
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.
- Returns:
The function modifies the file in place and does not return any values.
- Return type:
None
Note
The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.
Example
>>> _update_parquet_metadata( ... Path('data.pqt'), ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
File Operations
Utility functions for electrophysiological data processing and file management.
This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.
The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation
Functions
- setup_output_directory
Set up hierarchical output directory structure for probe and snippet data
- get_aggregated_snippets_df
Aggregate metadata from all snippets in a probe-level directory
- add_metadata_to_parquet_files
Add metadata attributes to all Parquet files in a snippet directory
- _update_parquet_metadata
Update metadata for a single Parquet file (internal helper function)
Examples
>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
... 'output_dir': '/data/output',
... 'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... 't_start': 300.0,
... 'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
... base_level_dir='/data/output',
... snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
... pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... t_start=300.0,
... duration=5.0
... )
Notes
This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.
See Also
pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations
- ephysatlas.utils.setup_output_directory(params)[source]
Set up the output directory structure for probe and snippet data.
This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.
- Parameters:
params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming
- Returns:
A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory
If output_dir is None, returns (None, None).
- Return type:
tuple[Path, Path]
- Raises:
ValueError – If neither pid nor filename is provided.
Note
The function creates a hierarchical structure:
Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration
Example output structure:
|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c | |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0 | `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0 `-- af0a0534-9cdc-4a29-93c0-1342891d74ec |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0 `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
- ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]
Get a dataframe of metadata info for all snippets in the probe level directory.
This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.
- Parameters:
probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.
- Returns:
- DataFrame containing metadata from all snippets, with each
row representing one snippet and columns representing different metadata attributes.
- Return type:
pd.DataFrame
Note
The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.
- ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]
Add metadata attributes to all Parquet files in a snippet-level directory.
This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.
- Parameters:
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.
- Returns:
The function modifies files in place and does not return any values.
- Return type:
None
Note
The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.
Example
>>> add_metadata_to_parquet_files( ... base_level_dir='/data/probe1', ... snippet_level_dir='snippet_001', ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
- ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]
Update metadata attributes for a single Parquet file.
This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.
- Parameters:
file_path (Path) – Path to the Parquet file to be updated.
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.
- Returns:
The function modifies the file in place and does not return any values.
- Return type:
None
Note
The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.
Example
>>> _update_parquet_metadata( ... Path('data.pqt'), ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
Internal Functions
Utility functions for electrophysiological data processing and file management.
This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.
The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation
Functions
- setup_output_directory
Set up hierarchical output directory structure for probe and snippet data
- get_aggregated_snippets_df
Aggregate metadata from all snippets in a probe-level directory
- add_metadata_to_parquet_files
Add metadata attributes to all Parquet files in a snippet directory
- _update_parquet_metadata
Update metadata for a single Parquet file (internal helper function)
Examples
>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
... 'output_dir': '/data/output',
... 'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... 't_start': 300.0,
... 'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
... base_level_dir='/data/output',
... snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
... pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... t_start=300.0,
... duration=5.0
... )
Notes
This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.
See Also
pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations
- ephysatlas.utils.setup_output_directory(params)[source]
Set up the output directory structure for probe and snippet data.
This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.
- Parameters:
params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming
- Returns:
A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory
If output_dir is None, returns (None, None).
- Return type:
tuple[Path, Path]
- Raises:
ValueError – If neither pid nor filename is provided.
Note
The function creates a hierarchical structure:
Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration
Example output structure:
|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c | |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0 | `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0 `-- af0a0534-9cdc-4a29-93c0-1342891d74ec |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0 `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
- ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]
Get a dataframe of metadata info for all snippets in the probe level directory.
This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.
- Parameters:
probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.
- Returns:
- DataFrame containing metadata from all snippets, with each
row representing one snippet and columns representing different metadata attributes.
- Return type:
pd.DataFrame
Note
The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.
- ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]
Add metadata attributes to all Parquet files in a snippet-level directory.
This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.
- Parameters:
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.
- Returns:
The function modifies files in place and does not return any values.
- Return type:
None
Note
The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.
Example
>>> add_metadata_to_parquet_files( ... base_level_dir='/data/probe1', ... snippet_level_dir='snippet_001', ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
- ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]
Update metadata attributes for a single Parquet file.
This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.
- Parameters:
file_path (Path) – Path to the Parquet file to be updated.
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.
- Returns:
The function modifies the file in place and does not return any values.
- Return type:
None
Note
The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.
Example
>>> _update_parquet_metadata( ... Path('data.pqt'), ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
Usage Examples
Utility functions for electrophysiological data processing and file management.
This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.
The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation
Functions
- setup_output_directory
Set up hierarchical output directory structure for probe and snippet data
- get_aggregated_snippets_df
Aggregate metadata from all snippets in a probe-level directory
- add_metadata_to_parquet_files
Add metadata attributes to all Parquet files in a snippet directory
- _update_parquet_metadata
Update metadata for a single Parquet file (internal helper function)
Examples
>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
... 'output_dir': '/data/output',
... 'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... 't_start': 300.0,
... 'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
... base_level_dir='/data/output',
... snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
... pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
... t_start=300.0,
... duration=5.0
... )
Notes
This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.
See Also
pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations
- ephysatlas.utils.setup_output_directory(params)[source]
Set up the output directory structure for probe and snippet data.
This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.
- Parameters:
params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming
- Returns:
A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory
If output_dir is None, returns (None, None).
- Return type:
tuple[Path, Path]
- Raises:
ValueError – If neither pid nor filename is provided.
Note
The function creates a hierarchical structure:
Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration
Example output structure:
|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c | |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0 | `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0 `-- af0a0534-9cdc-4a29-93c0-1342891d74ec |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0 `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
- ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]
Get a dataframe of metadata info for all snippets in the probe level directory.
This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.
- Parameters:
probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.
- Returns:
- DataFrame containing metadata from all snippets, with each
row representing one snippet and columns representing different metadata attributes.
- Return type:
pd.DataFrame
Note
The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.
- ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]
Add metadata attributes to all Parquet files in a snippet-level directory.
This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.
- Parameters:
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.
- Returns:
The function modifies files in place and does not return any values.
- Return type:
None
Note
The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.
Example
>>> add_metadata_to_parquet_files( ... base_level_dir='/data/probe1', ... snippet_level_dir='snippet_001', ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )
- ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]
Update metadata attributes for a single Parquet file.
This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.
- Parameters:
file_path (Path) – Path to the Parquet file to be updated.
**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.
- Returns:
The function modifies the file in place and does not return any values.
- Return type:
None
Note
The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.
Example
>>> _update_parquet_metadata( ... Path('data.pqt'), ... pid='probe1', ... t_start=100.5, ... duration=30.0 ... )