Skip to content

DataAccessor

Bases: BaseModel


              flowchart TD
              technologydata.data_accessor.DataAccessor[DataAccessor]

              

              click technologydata.data_accessor.DataAccessor href "" "technologydata.data_accessor.DataAccessor"
            

Access data from a versioned data source.

This class provides a standardized interface to locate and load technology datasets from predefined data sources. It can either load a specific version from the local storage, automatically determining and loading the latest available version, or download data from a remote URL.

Attributes:

  • data_source (str) –

    The name of the data source to access, as defined in the DataSourceName enumeration.

  • version ((str, optional)) –

    The specific version string of the data to load (e.g., "v1.0.0"). If not provided, the latest version will be automatically determined and used. Default is None.

  • data_path ((Path, optional)) –

    The path to the data source directory. If not provided, the parsers directory of the installed package is used (src/technologydata/parsers in a repository checkout).

Methods:

  • download –

    Download and load technology data from a remote URL.

  • ensure_path_exists –

    Ensure the provided data directory exists, creating it if necessary.

  • get_latest_version_string –

    Find the latest version string for the data source.

  • load –

    Load the default 'technologies.json' from the package data.

  • parse –

    Run the parser for the specified data source and version.

data_path class-attribute instance-attribute

data_path: Annotated[Path, Field(description='The base directory path where data sources are located.')] = path_parsers

data_source instance-attribute

data_source: Annotated[str, Field(description='The name of the data source.')]

version class-attribute instance-attribute

version: Annotated[str | None, Field(description='The version of the data source.')] = None

download

download(base_url: str, use_cache: bool = True) -> DataPackage

Download and load technology data from a remote URL.

This method downloads technologies.json and sources.json files from the specified base URL and loads them into a DataPackage instance. The sources.json file is optional; if not found, sources will be extracted from technologies.

By default, if the data is already cached locally, it will be loaded from the cache instead of re-downloading. Set use_cache=False to force a fresh download.

Important: The data_source and version attributes of the DataAccessor instance determine the local storage location for downloaded files. The base_url parameter determines what data is downloaded. It is the caller's responsibility to ensure these are consistent - the method does not validate that the URL content matches the specified version.

The downloaded files are stored locally in the directory structure: {data_path}/{data_source}/{version}/technologies.json and {data_path}/{data_source}/{version}/sources.json, where data_path defaults to src/technologydata/parsers/. This allows the data to be accessed later via the load() method without re-downloading.

Parameters:

  • base_url (str) –

    Base URL where the JSON files are hosted. The method will attempt to download technologies.json and sources.json from this location. The URL should point to the directory containing these files (without the filename itself). The caller must ensure this URL corresponds to the data_source and version specified in the DataAccessor instance.

  • use_cache (bool, default: True ) –

    If True (default), check for locally cached files first and use them if available. If False, always download from the URL even if cached files exist.

Returns:

  • DataPackage –

    An instance of DataPackage initialized with the downloaded data.

Raises:

  • HTTPError –

    If the HTTP request to download technologies.json or sources.json fails.

  • ConnectionError –

    If there is a network connectivity issue.

  • ValueError –

    If the version attribute is not set (required for determining storage location).

Notes
  • The data_source and version attributes must be set before calling this method.
  • The method does not validate that the downloaded data matches the specified version - this is the caller's responsibility.
  • The downloaded files persist on disk and can be reused in subsequent sessions by calling load() with the same data source and version.
  • When use_cache=True, only the existence of technologies.json is checked to determine if the cache is valid.

Examples:

Download v10 data and store it as v10 (uses cache if available):

>>> accessor = DataAccessor(data_source="dea_energy_storage", version="v10")
>>> base_url = "https://example.com/data/dea_energy_storage/v10/"
>>> dp = accessor.download(base_url)
>>> # If not cached, files are downloaded from the URL and stored at:
>>> # src/technologydata/parsers/dea_energy_storage/v10/

Force a fresh download, ignoring cache:

>>> dp = accessor.download(base_url, use_cache=False)

ensure_path_exists staticmethod

ensure_path_exists(input_data_path: Path) -> None

Ensure the provided data directory exists, creating it if necessary.

Creates the data directory and any parent directories as needed. If the directory already exists, no action is taken.

Parameters:

  • input_data_path (Path) –

    The base directory path where data sources are located.

Returns:

  • None –
Notes

This method uses mkdir(parents=True, exist_ok=True) to safely create the directory structure without raising an error if the directory already exists.

get_latest_version_string staticmethod

get_latest_version_string(data_source_path_list: list[Path]) -> str

Find the latest version string for the data source.

Returns:

  • str –

    The string of the latest version (e.g., 'v10', 'v1.0.0').

Raises:

  • FileNotFoundError –

    If the data source directory or valid version directories are not found.

load

load() -> DataPackage

Load the default 'technologies.json' from the package data.

Returns:

  • DataPackage –

    An instance of DataPackage initialized with the requested data.

Raises:

  • FileNotFoundError –

    If the data source directory or the specified version directory is not found.

  • ValueError: –

    If the specified version is not found in the data source directory.

parse

parse(input_file_names: list[str], num_digits: int = 4, archive_source: bool = False, filter_params: bool = False, export_schema: bool = False) -> None

Run the parser for the specified data source and version.

This method locates the appropriate parser for the given data source and version, and executes it to generate the technology data package.

Parameters:

  • input_file_names (list[str]) –

    The names of the input files within the data source directory in 'raw/{data_source}/'. Can be: - A single element list with single-file data sources (e.g., ["Technology_datasheet_for_energy_storage.xlsx"]) - A list of filenames for multi-file data sources (e.g., ["usa.csv", "other.csv"])

  • num_digits (int, default: 4 ) –

    Number of significant digits to round the values. Default is 4.

  • archive_source (bool, default: False ) –

    Store the source object on the Wayback Machine. Default is False.

  • filter_params (bool, default: False ) –

    Filter the parameters stored to technologies.json. Default is False.

  • export_schema (bool, default: False ) –

    Export the Source/TechnologyCollection schemas. Default is False.

Raises:

  • ValueError –

    If the specified data source or version is not supported.

  • FileNotFoundError –

    If the required input data file is not found.