Utils Module

The utils module provides utility functions for data management, file operations, and metadata handling.

Module Overview

This module includes: * Directory and output management utilities * Metadata handling for parquet files * Data aggregation and processing functions * File operation helpers

Core Functions

Directory Management

Utility functions for electrophysiological data processing and file management.

This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.

The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation

Functions

setup_output_directory

Set up hierarchical output directory structure for probe and snippet data

get_aggregated_snippets_df

Aggregate metadata from all snippets in a probe-level directory

add_metadata_to_parquet_files

Add metadata attributes to all Parquet files in a snippet directory

_update_parquet_metadata

Update metadata for a single Parquet file (internal helper function)

Examples

>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
...     'output_dir': '/data/output',
...     'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     't_start': 300.0,
...     'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/output',
...     snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
...     pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     t_start=300.0,
...     duration=5.0
... )

Notes

This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.

See Also

pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations

ephysatlas.utils.setup_output_directory(params)[source]

Set up the output directory structure for probe and snippet data.

This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.

Parameters:

params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming

Returns:

A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory

If output_dir is None, returns (None, None).

Return type:

tuple[Path, Path]

Raises:

ValueError – If neither pid nor filename is provided.

Note

The function creates a hierarchical structure:

Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration

Example output structure:

|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c
|   |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0
|   `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0
`-- af0a0534-9cdc-4a29-93c0-1342891d74ec
    |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0
    `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]

Get a dataframe of metadata info for all snippets in the probe level directory.

This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.

Parameters:

probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.

Returns:

DataFrame containing metadata from all snippets, with each

row representing one snippet and columns representing different metadata attributes.

Return type:

pd.DataFrame

Note

The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.

ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]

Add metadata attributes to all Parquet files in a snippet-level directory.

This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.

Parameters:

**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.

Returns:

The function modifies files in place and does not return any values.

Return type:

None

Note

The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.

Example

>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/probe1',
...     snippet_level_dir='snippet_001',
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )
ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]

Update metadata attributes for a single Parquet file.

This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.

Parameters:
  • file_path (Path) – Path to the Parquet file to be updated.

  • **snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.

Returns:

The function modifies the file in place and does not return any values.

Return type:

None

Note

The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.

Example

>>> _update_parquet_metadata(
...     Path('data.pqt'),
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )

Metadata Management

Utility functions for electrophysiological data processing and file management.

This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.

The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation

Functions

setup_output_directory

Set up hierarchical output directory structure for probe and snippet data

get_aggregated_snippets_df

Aggregate metadata from all snippets in a probe-level directory

add_metadata_to_parquet_files

Add metadata attributes to all Parquet files in a snippet directory

_update_parquet_metadata

Update metadata for a single Parquet file (internal helper function)

Examples

>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
...     'output_dir': '/data/output',
...     'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     't_start': 300.0,
...     'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/output',
...     snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
...     pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     t_start=300.0,
...     duration=5.0
... )

Notes

This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.

See Also

pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations

ephysatlas.utils.setup_output_directory(params)[source]

Set up the output directory structure for probe and snippet data.

This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.

Parameters:

params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming

Returns:

A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory

If output_dir is None, returns (None, None).

Return type:

tuple[Path, Path]

Raises:

ValueError – If neither pid nor filename is provided.

Note

The function creates a hierarchical structure:

Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration

Example output structure:

|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c
|   |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0
|   `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0
`-- af0a0534-9cdc-4a29-93c0-1342891d74ec
    |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0
    `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]

Get a dataframe of metadata info for all snippets in the probe level directory.

This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.

Parameters:

probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.

Returns:

DataFrame containing metadata from all snippets, with each

row representing one snippet and columns representing different metadata attributes.

Return type:

pd.DataFrame

Note

The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.

ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]

Add metadata attributes to all Parquet files in a snippet-level directory.

This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.

Parameters:

**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.

Returns:

The function modifies files in place and does not return any values.

Return type:

None

Note

The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.

Example

>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/probe1',
...     snippet_level_dir='snippet_001',
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )
ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]

Update metadata attributes for a single Parquet file.

This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.

Parameters:
  • file_path (Path) – Path to the Parquet file to be updated.

  • **snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.

Returns:

The function modifies the file in place and does not return any values.

Return type:

None

Note

The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.

Example

>>> _update_parquet_metadata(
...     Path('data.pqt'),
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )

File Operations

Utility functions for electrophysiological data processing and file management.

This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.

The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation

Functions

setup_output_directory

Set up hierarchical output directory structure for probe and snippet data

get_aggregated_snippets_df

Aggregate metadata from all snippets in a probe-level directory

add_metadata_to_parquet_files

Add metadata attributes to all Parquet files in a snippet directory

_update_parquet_metadata

Update metadata for a single Parquet file (internal helper function)

Examples

>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
...     'output_dir': '/data/output',
...     'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     't_start': 300.0,
...     'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/output',
...     snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
...     pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     t_start=300.0,
...     duration=5.0
... )

Notes

This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.

See Also

pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations

ephysatlas.utils.setup_output_directory(params)[source]

Set up the output directory structure for probe and snippet data.

This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.

Parameters:

params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming

Returns:

A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory

If output_dir is None, returns (None, None).

Return type:

tuple[Path, Path]

Raises:

ValueError – If neither pid nor filename is provided.

Note

The function creates a hierarchical structure:

Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration

Example output structure:

|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c
|   |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0
|   `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0
`-- af0a0534-9cdc-4a29-93c0-1342891d74ec
    |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0
    `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]

Get a dataframe of metadata info for all snippets in the probe level directory.

This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.

Parameters:

probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.

Returns:

DataFrame containing metadata from all snippets, with each

row representing one snippet and columns representing different metadata attributes.

Return type:

pd.DataFrame

Note

The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.

ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]

Add metadata attributes to all Parquet files in a snippet-level directory.

This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.

Parameters:

**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.

Returns:

The function modifies files in place and does not return any values.

Return type:

None

Note

The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.

Example

>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/probe1',
...     snippet_level_dir='snippet_001',
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )
ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]

Update metadata attributes for a single Parquet file.

This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.

Parameters:
  • file_path (Path) – Path to the Parquet file to be updated.

  • **snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.

Returns:

The function modifies the file in place and does not return any values.

Return type:

None

Note

The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.

Example

>>> _update_parquet_metadata(
...     Path('data.pqt'),
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )

Internal Functions

Utility functions for electrophysiological data processing and file management.

This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.

The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation

Functions

setup_output_directory

Set up hierarchical output directory structure for probe and snippet data

get_aggregated_snippets_df

Aggregate metadata from all snippets in a probe-level directory

add_metadata_to_parquet_files

Add metadata attributes to all Parquet files in a snippet directory

_update_parquet_metadata

Update metadata for a single Parquet file (internal helper function)

Examples

>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
...     'output_dir': '/data/output',
...     'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     't_start': 300.0,
...     'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/output',
...     snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
...     pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     t_start=300.0,
...     duration=5.0
... )

Notes

This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.

See Also

pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations

ephysatlas.utils.setup_output_directory(params)[source]

Set up the output directory structure for probe and snippet data.

This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.

Parameters:

params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming

Returns:

A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory

If output_dir is None, returns (None, None).

Return type:

tuple[Path, Path]

Raises:

ValueError – If neither pid nor filename is provided.

Note

The function creates a hierarchical structure:

Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration

Example output structure:

|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c
|   |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0
|   `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0
`-- af0a0534-9cdc-4a29-93c0-1342891d74ec
    |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0
    `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]

Get a dataframe of metadata info for all snippets in the probe level directory.

This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.

Parameters:

probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.

Returns:

DataFrame containing metadata from all snippets, with each

row representing one snippet and columns representing different metadata attributes.

Return type:

pd.DataFrame

Note

The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.

ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]

Add metadata attributes to all Parquet files in a snippet-level directory.

This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.

Parameters:

**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.

Returns:

The function modifies files in place and does not return any values.

Return type:

None

Note

The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.

Example

>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/probe1',
...     snippet_level_dir='snippet_001',
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )
ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]

Update metadata attributes for a single Parquet file.

This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.

Parameters:
  • file_path (Path) – Path to the Parquet file to be updated.

  • **snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.

Returns:

The function modifies the file in place and does not return any values.

Return type:

None

Note

The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.

Example

>>> _update_parquet_metadata(
...     Path('data.pqt'),
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )

Usage Examples

Utility functions for electrophysiological data processing and file management.

This module provides utility functions for managing output directories, handling metadata, and working with Parquet files in electrophysiological data analysis pipelines. It includes functions for directory structure creation, metadata aggregation, and file attribute management.

The module includes: - Output directory structure setup and management - Metadata aggregation from multiple snippet directories - Parquet file metadata management and updates - File path handling and validation

Functions

setup_output_directory

Set up hierarchical output directory structure for probe and snippet data

get_aggregated_snippets_df

Aggregate metadata from all snippets in a probe-level directory

add_metadata_to_parquet_files

Add metadata attributes to all Parquet files in a snippet directory

_update_parquet_metadata

Update metadata for a single Parquet file (internal helper function)

Examples

>>> from ephysatlas.utils import setup_output_directory, add_metadata_to_parquet_files
>>> from pathlib import Path
>>>
>>> # Set up output directory structure
>>> params = {
...     'output_dir': '/data/output',
...     'pid': '76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     't_start': 300.0,
...     'duration': 5.0
... }
>>> probe_dir, snippet_dir = setup_output_directory(params)
>>>
>>> # Add metadata to Parquet files
>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/output',
...     snippet_level_dir='probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0',
...     pid='76ed566f-59dd-47ff-8ba7-59b11d09b67c',
...     t_start=300.0,
...     duration=5.0
... )

Notes

This module handles the creation of hierarchical directory structures for organizing electrophysiological data by probe ID and time snippets. It supports both .parquet and .pqt file extensions and provides robust metadata management capabilities for tracking data provenance and processing parameters.

See Also

pandas.DataFrame.attrs : DataFrame attributes for metadata storage pathlib.Path : Path manipulation and directory operations

ephysatlas.utils.setup_output_directory(params)[source]

Set up the output directory structure for probe and snippet data.

This function creates a hierarchical directory structure for organizing electrophysiological data by probe ID and time snippets. It supports both probe ID-based and file-based directory naming.

Parameters:

params (Dict[str, Any]) – Dictionary containing configuration parameters. Must include: - output_dir (str, optional): Base output directory path - pid (str, optional): Probe ID for directory naming - filename (str, optional): AP file path for hash-based naming - t_start (float): Start time for snippet naming - duration (float): Duration for snippet naming

Returns:

A tuple containing: - probe_level_dir (Path): Path to the probe-level directory - snippet_level_dir (Path): Path to the snippet-level directory

If output_dir is None, returns (None, None).

Return type:

tuple[Path, Path]

Raises:

ValueError – If neither pid nor filename is provided.

Note

The function creates a hierarchical structure:

Probe level: Uses pid or hash of filename Snippet level: Uses probe info, t_start, and duration

Example output structure:

|-- 76ed566f-59dd-47ff-8ba7-59b11d09b67c
|   |-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_000300.0_05.0
|   `-- probe_76ed566f-59dd-47ff-8ba7-59b11d09b67c_003000.0_05.0
`-- af0a0534-9cdc-4a29-93c0-1342891d74ec
    |-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_000300.0_05.0
    `-- probe_af0a0534-9cdc-4a29-93c0-1342891d74ec_003000.0_05.0
ephysatlas.utils.get_aggregated_snippets_df(probe_level_dir)[source]

Get a dataframe of metadata info for all snippets in the probe level directory.

This function scans a probe-level directory and aggregates metadata from all snippet subdirectories. It reads metadata from Parquet files (.parquet and .pqt) and combines them into a single DataFrame for analysis.

Parameters:

probe_level_dir (Path) – Path to the probe-level directory containing snippet subdirectories.

Returns:

DataFrame containing metadata from all snippets, with each

row representing one snippet and columns representing different metadata attributes.

Return type:

pd.DataFrame

Note

The function looks for both .parquet and .pqt file extensions in each snippet directory. It extracts metadata from the DataFrame’s attrs dictionary and combines them across all snippets.

ephysatlas.utils.add_metadata_to_parquet_files(**snippet_attrs)[source]

Add metadata attributes to all Parquet files in a snippet-level directory.

This function takes snippet attributes and adds them as metadata to all .parquet and .pqt files found in the specified snippet directory. The metadata is useful for tracking provenance, parameters, and other contextual information.

Parameters:

**snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata to add to the Parquet files. Must include: - base_level_dir (str): Base directory path - snippet_level_dir (str): Snippet directory name Additional key-value pairs will be added as metadata attributes.

Returns:

The function modifies files in place and does not return any values.

Return type:

None

Note

The function constructs the full snippet directory path from base_level_dir and snippet_level_dir. Both .parquet and .pqt file extensions are supported. If the directory doesn’t exist, a warning is logged but no error is raised. Each file is processed individually using _update_parquet_metadata.

Example

>>> add_metadata_to_parquet_files(
...     base_level_dir='/data/probe1',
...     snippet_level_dir='snippet_001',
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )
ephysatlas.utils._update_parquet_metadata(file_path, **snippet_attrs)[source]

Update metadata attributes for a single Parquet file.

This helper function reads a Parquet file, adds the provided metadata attributes to the DataFrame’s attrs dictionary, and writes the file back to disk.

Parameters:
  • file_path (Path) – Path to the Parquet file to be updated.

  • **snippet_attrs (Dict[str, Any]) – Keyword arguments containing metadata attributes to add to the file. These will be stored in the DataFrame’s attrs dictionary.

Returns:

The function modifies the file in place and does not return any values.

Return type:

None

Note

The function reads the entire Parquet file into memory, modifies it, and writes it back. All provided snippet_attrs are added to the DataFrame’s attrs dictionary. If an error occurs during processing, it is logged as a warning but doesn’t stop execution. This function is designed to be called by add_metadata_to_parquet_files for batch processing.

Example

>>> _update_parquet_metadata(
...     Path('data.pqt'),
...     pid='probe1',
...     t_start=100.5,
...     duration=30.0
... )