{ "cells": [ { "cell_type": "markdown", "id": "c3056ef8", "metadata": {}, "source": [ "# Getting started\n", "\n", "The epidatpy package provides access to all the endpoints of the [Delphi Epidata\n", "API](https://cmu-delphi.github.io/delphi-epidata/), and can be used to make\n", "requests for specific signals on specific dates and in select geographic\n", "regions. It is widely used in epidemiological research, real-time forecasting\n", "models, and public health dashboards." ] }, { "cell_type": "markdown", "id": "1c158c4c", "metadata": {}, "source": [ "## Setup\n", "\n", "### Installation\n", "\n", "Install the stable version from PyPI:\n", "\n", "```sh\n", "pip install epidatpy\n", "```\n", "\n", "Or, for the development version, install from GitHub:\n", "\n", "```sh\n", "pip install \"git+https://github.com/cmu-delphi/epidatpy.git#egg=epidatpy\"\n", "```\n", "\n", "### API keys\n", "\n", "The Delphi API requires a (free) API key for full functionality. While most\n", "endpoints are available without one, there are\n", "[limits on API usage for anonymous users](https://cmu-delphi.github.io/delphi-epidata/api/api_keys.html),\n", "including a rate limit.\n", "\n", "To generate your key,\n", "[register for a pseudo-anonymous account](https://api.delphi.cmu.edu/epidata/admin/registration_form).\n", "`epidatpy` reads the key from the `DELPHI_EPIDATA_KEY` environment variable. We\n", "recommend storing it in a `.env` file, loading it with\n", "[python-dotenv](https://github.com/theskumar/python-dotenv), and adding `.env`\n", "to your `.gitignore`." ] }, { "cell_type": "markdown", "id": "05dcdb96", "metadata": {}, "source": [ "## The Delphi V5 API\n", "\n", "`epidatpy` allows three categories of data access to the Delphi V5 API:\n", "\n", "- `epidata_snapshot()` provides a specific view of how a dataset looked at a\n", " point in time.\n", "- `epidata_archive()` fetches all versions of a dataset across time,\n", " representing the full revision history.\n", "- `epidata_aux()` accesses source-specific auxiliary tables containing\n", " metadata, laboratory protocols, or additional static keys (such as NWSS\n", " wastewater facility descriptions).\n", "\n", "Additionally, `epidata_meta()` provides access to system metadata to list\n", "available sources, signals, geographic granularities, and date ranges.\n", "`epidata()` is a convenience wrapper that routes to `epidata_archive()` if you\n", "pass `report_time`, or to `epidata_snapshot()` if you pass `snapshot_date` (or\n", "neither).\n", "\n", "The older V4 (`pub_covidcast`) and V3 (`pub_fluview`, `pub_flusurv`, ...)\n", "endpoints still work, but starting in October 2026 they are tentatively\n", "deprecated in favor of V5. New code should start on V5; see the\n", "[migration guide](migration_guide.ipynb) for how to move existing code." ] }, { "cell_type": "markdown", "id": "6cd48851", "metadata": {}, "source": [ "## Basic usage\n", "\n", "To make a request of a particular data source at a specific point in time, we\n", "use `epidata_snapshot()`. This function needs the source name, signal name, and a\n", "geographic level in order to complete a query.\n", "\n", "Suppose we are interested in the `nssp` source, which provides access to a\n", "[wide range of](https://cmu-delphi.github.io/delphi-epidata/api/v5-signals/nssp.html)\n", "emergency department visits data. `epidata_meta()` tells us which signals and\n", "geographic levels the source offers, and the range of dates available:" ] }, { "cell_type": "code", "execution_count": null, "id": "83f16458", "metadata": {}, "outputs": [], "source": [ "import pandas as pd\n", "\n", "pd.set_option(\"display.max_columns\", None)\n", "pd.set_option(\"display.max_rows\", 10)\n", "pd.set_option(\"display.width\", 1000)" ] }, { "cell_type": "code", "execution_count": null, "id": "837e6cf3", "metadata": {}, "outputs": [], "source": [ "from epidatpy import EpiDataContext, EpiRange\n", "\n", "epidata = EpiDataContext()\n", "\n", "meta = epidata.epidata_meta(source=\"nssp\")\n", "print(meta[\"signals\"])\n", "print(meta[\"geo_types\"])\n", "print(meta[\"reference_time_range\"])\n", "print(meta[\"report_time_range\"])" ] }, { "cell_type": "markdown", "id": "842bcc2d", "metadata": {}, "source": [ "All of the fetch functions return an `EpiDataCall`, a not-yet-executed query\n", "that you can inspect (`.request_url()` shows the underlying API request) and\n", "then execute with `.df()` to obtain a pandas DataFrame:" ] }, { "cell_type": "code", "execution_count": null, "id": "a8e6ef34", "metadata": {}, "outputs": [], "source": [ "# Obtain the latest snapshot of the influenza ED-visit percentage\n", "# from the NSSP source for the US\n", "apicall = epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"nation\",\n", ")\n", "print(apicall)\n", "\n", "us_flu = apicall.df()\n", "us_flu" ] }, { "cell_type": "markdown", "id": "cd938f6f", "metadata": {}, "source": [ "Each row represents one observation for the US on one date. The location is\n", "given in the `geo_value` column, the date it describes in the `reference_time`\n", "column, the value of the requested signal in `value`, and the publication time\n", "in `report_time` (a UTC timestamp).\n", "\n", "The Delphi V5 API makes signals available at different geographic levels,\n", "depending on the source. To request signals for all states instead of the\n", "entire US, we change the `geo_type` argument. This automatically returns all\n", "available data for that geo type:" ] }, { "cell_type": "code", "execution_count": null, "id": "c66a36d0", "metadata": {}, "outputs": [], "source": [ "# Obtain the latest snapshot of the influenza ED-visit percentage\n", "# from the NSSP source across all available dates and states\n", "epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", ").df()" ] }, { "cell_type": "markdown", "id": "35c96d65", "metadata": {}, "source": [ "You can query multiple signals in a single request by passing a list to\n", "`signals`, and narrow the result with `geo_values` and a `reference_time` range\n", "(both are filtered locally after the request):" ] }, { "cell_type": "code", "execution_count": null, "id": "3c630250", "metadata": {}, "outputs": [], "source": [ "# Obtain both influenza and COVID-19 ED-visit percentages in a single query\n", "epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=[\"pct_ed_visits_influenza\", \"pct_ed_visits_covid\"],\n", " geo_type=\"state\",\n", " geo_values=\"pa\",\n", " reference_time=EpiRange(\"2024-12-01\", \"2024-12-15\"),\n", ").df()" ] }, { "cell_type": "markdown", "id": "31b71583", "metadata": {}, "source": [ "Alternatively, we can fetch the time series for a subset of states and reference\n", "dates by listing out the desired locations in `geo_values` and using a range in\n", "`reference_time`:" ] }, { "cell_type": "code", "execution_count": null, "id": "b79c9542", "metadata": {}, "outputs": [], "source": [ "# Obtain the data from January 1st, 2024 to January 1st, 2025\n", "# of the influenza ED-visit percentage from the NSSP source for\n", "# Pennsylvania, California, and Florida\n", "states_flu = epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", " geo_values=[\"pa\", \"ca\", \"fl\"],\n", " reference_time=EpiRange(\"2024-01-01\", \"2025-01-01\"),\n", ").df()\n", "states_flu" ] }, { "cell_type": "markdown", "id": "b8a5a516", "metadata": {}, "source": [ "`geo_type` also accepts several values. The server handles one geographic level\n", "per request, so `epidatpy` issues one request per level and concatenates the\n", "results:" ] }, { "cell_type": "code", "execution_count": null, "id": "644bd11b", "metadata": {}, "outputs": [], "source": [ "epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=[\"nation\", \"hhs\"],\n", " reference_time=\"2024-12-07\",\n", ").df()" ] }, { "cell_type": "markdown", "id": "85413583", "metadata": {}, "source": [ "## Getting versioned data\n", "\n", "The Delphi V5 API stores a historical record of all data, including corrections\n", "and updates, which is particularly useful for accurately backtesting forecasting\n", "models. To retrieve versioned data in `epidata_snapshot()`, we use the\n", "`snapshot_date` argument, which fetches the data as it was known on a specific\n", "date (a `date`, a `YYYY-MM-DD` string, or a `datetime` for an exact instant):" ] }, { "cell_type": "code", "execution_count": null, "id": "3bbb494a", "metadata": {}, "outputs": [], "source": [ "# Obtain the influenza ED-visit percentage from NSSP for Pennsylvania\n", "# as it was known on 2025-01-01\n", "epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", " geo_values=\"pa\",\n", " snapshot_date=\"2025-01-01\",\n", ").df()" ] }, { "cell_type": "markdown", "id": "ff16267f", "metadata": {}, "source": [ "To request all versions of the data issued within a specific time range, we use\n", "`epidata_archive()` with the `report_time` argument. It accepts a comparison\n", "string (such as `\"<2025-01-15\"` or `\">=2024-12-01\"`) or an `EpiRange` for an\n", "inclusive range of report dates. A bare date is not accepted: for the data as it\n", "looked on a single date, use `epidata_snapshot()` instead." ] }, { "cell_type": "code", "execution_count": null, "id": "1cb9b491", "metadata": {}, "outputs": [], "source": [ "# See how the estimate for a SINGLE reference date (2024-12-07) evolved\n", "# by fetching all reports issued in December 2024 and early January 2025\n", "epidata.epidata_archive(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", " geo_values=\"pa\",\n", " reference_time=\"2024-12-07\",\n", " report_time=EpiRange(\"2024-12-01\", \"2025-01-15\"),\n", ").df()" ] }, { "cell_type": "code", "execution_count": null, "id": "a9965e83", "metadata": {}, "outputs": [], "source": [ "# Everything reported strictly before 2024-12-15 for the same reference date\n", "epidata.epidata_archive(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", " geo_values=\"pa\",\n", " reference_time=\"2024-12-07\",\n", " report_time=\"<2024-12-15\",\n", ").df()" ] }, { "cell_type": "markdown", "id": "9be1ba72", "metadata": {}, "source": [ "See the [versioned data](versioned_data.ipynb) notebook for details and more\n", "ways to specify versioned data." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Auxiliary data\n", "\n", "Some sources carry extra columns alongside the signal data, such as the\n", "population served by each NWSS sewershed or its site metadata.\n", "`epidata_aux()` retrieves that auxiliary data, either on its own or merged\n", "onto a signal pull.\n", "\n", "To pull it directly by source, pass named filters on the source's key\n", "columns as keyword arguments; `columns` selects specific fields. To see\n", "which key columns a source has, consult its page in the [V5 signals\n", "documentation](https://cmu-delphi.github.io/delphi-epidata/api/v5_signals.html)." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "aux_data = epidata.epidata_aux(\n", " source=\"nwss\",\n", " pcr_target=\"sars-cov-2\",\n", " sample_index=[\"92012\", \"92013\"],\n", ").df()\n", "aux_data.head()" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "You can also attach the auxiliary columns to a signal pull by passing the\n", "result of `epidata_snapshot()` or `epidata_archive()` straight to\n", "`epidata_aux()`. It fetches the matching auxiliary data and left-joins it on\n", "the shared key columns; the key filters are inferred from the base dataset,\n", "so you don't have to repeat them." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Fetch signal data for a specific sewershed\n", "nwss_data = epidata.epidata_snapshot(\n", " source=\"nwss\",\n", " signals=\"covid_avg_conc\",\n", " geo_type=\"sewershed\",\n", " geo_values=\"128\",\n", " reference_time=EpiRange(\"2024-12-01\", \"2025-01-01\"),\n", ").df()\n", "nwss_data.head()" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# Attach auxiliary metadata\n", "nwss_merged = epidata.epidata_aux(nwss_data)\n", "nwss_merged.head()" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Advanced queries\n", "\n", "### Dry runs\n", "\n", "To inspect the API request a query would make without fetching anything,\n", "build the call and read its `request_url()`. This works for\n", "`epidata_snapshot()`, `epidata_archive()`, and `epidata_aux()`, since each\n", "returns an `EpiDataCall` that only contacts the server once you call `.df()`." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "dry_run_call = epidata.epidata_snapshot(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", ")\n", "dry_run_call.request_url()" ] }, { "cell_type": "markdown", "id": "71171fd2", "metadata": {}, "source": [ "## Plotting\n", "\n", "Because the output data is a standard pandas DataFrame, we can easily plot it\n", "using any of the available Python libraries:" ] }, { "cell_type": "code", "execution_count": null, "id": "5e7b8257", "metadata": {}, "outputs": [], "source": [ "import matplotlib.pyplot as plt\n", "\n", "plt.rcParams[\"figure.dpi\"] = 150\n", "\n", "fig, ax = plt.subplots(figsize=(8, 4))\n", "(\n", " states_flu.pivot_table(values=\"value\", index=\"reference_time\", columns=\"geo_value\").plot(\n", " xlabel=\"Date\", ylabel=\"% of ED visits\", ax=ax, linewidth=1.5\n", " )\n", ")\n", "ax.set_title(\"Influenza ED visits from NSSP, 2024\")\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "f3ce962e", "metadata": {}, "source": [ "### Plotting revision histories\n", "\n", "We can also visualize revision histories from `epidata_archive()`. Each line\n", "shows what the time series looked like as of a different publication date:" ] }, { "cell_type": "code", "execution_count": null, "id": "673f62db", "metadata": {}, "outputs": [], "source": [ "# Fetch revision history for Pennsylvania influenza ED visits\n", "pa_revisions = epidata.epidata_archive(\n", " source=\"nssp\",\n", " signals=\"pct_ed_visits_influenza\",\n", " geo_type=\"state\",\n", " geo_values=\"pa\",\n", " reference_time=EpiRange(\"2024-10-01\", \"2024-12-01\"),\n", " report_time=EpiRange(\"2024-11-01\", \"2025-01-01\"),\n", ").df()\n", "\n", "fig, ax = plt.subplots(figsize=(8, 4))\n", "for report_time, group in pa_revisions.groupby(\"report_time\"):\n", " group.sort_values(\"reference_time\").plot(\n", " x=\"reference_time\", y=\"value\", ax=ax, label=report_time.strftime(\"%Y-%m-%d\"), linewidth=1\n", " )\n", "ax.set_title(\"Revisions of NSSP influenza ED visits in Pennsylvania\")\n", "ax.set_xlabel(\"Observation date\")\n", "ax.set_ylabel(\"% of ED visits\")\n", "ax.legend(title=\"Report date\", fontsize=7, ncol=2)\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "51ff5f1e", "metadata": {}, "source": [ "## Available data sources and endpoints\n", "\n", "`epidatpy` provides access to a broad ecosystem of epidemiological data streams:\n", "\n", "- **V5 sources** provide access to active surveillance data queried via\n", " `epidata_snapshot()` and `epidata_archive()`. Discover them programmatically\n", " with `epidata_meta()` (no `source` argument lists every source) or\n", " interactively on the [Delphi EpiPortal](https://delphi.cmu.edu/epiportal/).\n", "- **Migrating endpoints** are legacy endpoints (such as `pub_covidcast()`,\n", " `pub_fluview()`, `pub_flusurv()`, and `pub_meta()`) transitioning to V5. See\n", " the [migration guide](migration_guide.ipynb) for argument mappings.\n", "- **Historical endpoints** provide access to datasets whose collection has\n", " ended (such as Google Flu Trends, Wikipedia article views, and historical\n", " hospitalization series), kept for retrospective analysis via `pub_*`\n", " functions.\n", " - **International endpoints** are a subset of historical datasets that focus\n", " on surveillance outside the United States (e.g., PAHO dengue with\n", " `pub_paho_dengue()` and ECDC ILI with `pub_ecdc_ili()`).\n", " - **Private endpoints** are restricted streams (e.g., CDC web metrics with\n", " `pvt_cdc()` and digital sensors with `pvt_sensors()`) that require dedicated\n", " secret authentication keys." ] }, { "cell_type": "code", "execution_count": null, "id": "05f47b15", "metadata": {}, "outputs": [], "source": [ "all_meta = epidata.epidata_meta()\n", "print(sorted(all_meta))" ] }, { "cell_type": "code", "execution_count": null, "id": "78d8a7fc", "metadata": {}, "outputs": [], "source": [ "from epidatpy import available_endpoints\n", "\n", "available_endpoints()" ] }, { "cell_type": "markdown", "id": "3afc61d9", "metadata": {}, "source": [ "See the [signal discovery](signal_discovery.ipynb) notebook for an in-depth\n", "guide to discovering signals, browsing metadata, and querying datasets across\n", "all these categories." ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python" } }, "nbformat": 4, "nbformat_minor": 5 }