DataAccessor
¶
Bases: BaseModel
flowchart TD
technologydata.data_accessor.DataAccessor[DataAccessor]
click technologydata.data_accessor.DataAccessor href "" "technologydata.data_accessor.DataAccessor"
Access data from a versioned data source.
This class provides a standardized interface to locate and load technology datasets from predefined data sources. It can either load a specific version from the local storage, automatically determining and loading the latest available version, or download data from a remote URL.
Attributes:
-
data_source(str) –The name of the data source to access, as defined in the
DataSourceNameenumeration. -
version((str, optional)) –The specific version string of the data to load (e.g., "v1.0.0"). If not provided, the latest version will be automatically determined and used. Default is None.
-
data_path((Path, optional)) –The path to the data source directory. If not provided, the
parsersdirectory of the installed package is used (src/technologydata/parsersin a repository checkout).
Methods:
-
download–Download and load technology data from a remote URL.
-
ensure_path_exists–Ensure the provided data directory exists, creating it if necessary.
-
get_latest_version_string–Find the latest version string for the data source.
-
load–Load the default 'technologies.json' from the package data.
-
parse–Run the parser for the specified data source and version.
data_path
class-attribute
instance-attribute
¶
data_path: Annotated[Path, Field(description='The base directory path where data sources are located.')] = path_parsers
data_source
instance-attribute
¶
data_source: Annotated[str, Field(description='The name of the data source.')]
version
class-attribute
instance-attribute
¶
version: Annotated[str | None, Field(description='The version of the data source.')] = None
download
¶
download(base_url: str, use_cache: bool = True) -> DataPackage
Download and load technology data from a remote URL.
This method downloads technologies.json and sources.json files from the specified base URL and loads them into a DataPackage instance. The sources.json file is optional; if not found, sources will be extracted from technologies.
By default, if the data is already cached locally, it will be loaded from the
cache instead of re-downloading. Set use_cache=False to force a fresh download.
Important: The data_source and version attributes of the
DataAccessor instance determine the local storage location for downloaded
files. The base_url parameter determines what data is downloaded.
It is the caller's responsibility to ensure these are consistent - the method
does not validate that the URL content matches the specified version.
The downloaded files are stored locally in the directory structure:
{data_path}/{data_source}/{version}/technologies.json and
{data_path}/{data_source}/{version}/sources.json, where data_path
defaults to src/technologydata/parsers/. This allows the data to be
accessed later via the load() method without re-downloading.
Parameters:
-
base_url(str) –Base URL where the JSON files are hosted. The method will attempt to download technologies.json and sources.json from this location. The URL should point to the directory containing these files (without the filename itself). The caller must ensure this URL corresponds to the data_source and version specified in the DataAccessor instance.
-
use_cache(bool, default:True) –If True (default), check for locally cached files first and use them if available. If False, always download from the URL even if cached files exist.
Returns:
-
DataPackage–An instance of DataPackage initialized with the downloaded data.
Raises:
-
HTTPError–If the HTTP request to download technologies.json or sources.json fails.
-
ConnectionError–If there is a network connectivity issue.
-
ValueError–If the version attribute is not set (required for determining storage location).
Notes
- The
data_sourceandversionattributes must be set before calling this method. - The method does not validate that the downloaded data matches the specified version - this is the caller's responsibility.
- The downloaded files persist on disk and can be reused in subsequent sessions
by calling
load()with the same data source and version. - When
use_cache=True, only the existence of technologies.json is checked to determine if the cache is valid.
Examples:
Download v10 data and store it as v10 (uses cache if available):
>>> accessor = DataAccessor(data_source="dea_energy_storage", version="v10")
>>> base_url = "https://example.com/data/dea_energy_storage/v10/"
>>> dp = accessor.download(base_url)
>>> # If not cached, files are downloaded from the URL and stored at:
>>> # src/technologydata/parsers/dea_energy_storage/v10/
Force a fresh download, ignoring cache:
>>> dp = accessor.download(base_url, use_cache=False)
ensure_path_exists
staticmethod
¶
ensure_path_exists(input_data_path: Path) -> None
Ensure the provided data directory exists, creating it if necessary.
Creates the data directory and any parent directories as needed. If the directory already exists, no action is taken.
Parameters:
-
input_data_path(Path) –The base directory path where data sources are located.
Returns:
-
None–
Notes
This method uses mkdir(parents=True, exist_ok=True) to safely create
the directory structure without raising an error if the directory
already exists.
get_latest_version_string
staticmethod
¶
get_latest_version_string(data_source_path_list: list[Path]) -> str
Find the latest version string for the data source.
Returns:
-
str–The string of the latest version (e.g., 'v10', 'v1.0.0').
Raises:
-
FileNotFoundError–If the data source directory or valid version directories are not found.
load
¶
load() -> DataPackage
Load the default 'technologies.json' from the package data.
Returns:
-
DataPackage–An instance of DataPackage initialized with the requested data.
Raises:
-
FileNotFoundError–If the data source directory or the specified version directory is not found.
-
ValueError:–If the specified version is not found in the data source directory.
parse
¶
parse(input_file_names: list[str], num_digits: int = 4, archive_source: bool = False, filter_params: bool = False, export_schema: bool = False) -> None
Run the parser for the specified data source and version.
This method locates the appropriate parser for the given data source and version, and executes it to generate the technology data package.
Parameters:
-
input_file_names(list[str]) –The names of the input files within the data source directory in 'raw/{data_source}/'. Can be: - A single element list with single-file data sources (e.g., ["Technology_datasheet_for_energy_storage.xlsx"]) - A list of filenames for multi-file data sources (e.g., ["usa.csv", "other.csv"])
-
num_digits(int, default:4) –Number of significant digits to round the values. Default is 4.
-
archive_source(bool, default:False) –Store the source object on the Wayback Machine. Default is False.
-
filter_params(bool, default:False) –Filter the parameters stored to technologies.json. Default is False.
-
export_schema(bool, default:False) –Export the Source/TechnologyCollection schemas. Default is False.
Raises:
-
ValueError–If the specified data source or version is not supported.
-
FileNotFoundError–If the required input data file is not found.