Data

This section of the webpage briefly outlines the datasets that should be used during the Hackathon and provides related resources and documentation. The Hackathon’s problem statements can be found here and should be referred to regarding the event’s purpose and objectives.

The human mobility datasets provided for the Hackathon fall into one of two categories: digital data (from Meta) or key informant survey data (from IOM). The table below outlines these, summarising their accessibility, processing and stability. In addition, contextual data on Pakistan is provided to help with analysis, as well as a variety of spatial boundaries that can be used to map the provided datasets. Each dataset is outlined below, alongside any available documentation and guidance for analysis.

Data Provider Spatial Detail Source Access Processing Stability
Facebook Population and Movement During Crisis Meta

800m tiles

Global admin 2

Meta apps (mainly Facebook) Free to researchers Computational Made available after a disaster event for 90 days
Community Needs Identification Data IOM

Province

District

Tehsil

Village/settlement

Key informants Published online (collection is time-consuming and costly) Manpower Ad Hoc survey rounds and reports after the disaster event

1. Data Access

All the data mentioned in this document is accessible via Zenodo for Hackathon participants.

Note

To access the Zenodo repository, you will need to sign up for the Hackathon and complete the data access form provided upon registration.

2. Datasets

2.1 Facebook Population and Movement During Crisis

  • Meta provides these data as part of their Data for Good programme in the aftermath of humanitarian disasters. Once uploaded, these data are available to researchers and policymakers for 90 days before being removed from the Data for Good platform. The data shows the number of Facebook users located within a given spatial unit at a given time.

  • For the Hackathon, we are using the data available in the immediate aftermath of flood events in Pakistan in 2022 and 2025. The data for the 2022 flood event has coverage of the vast majority of the area administered by Pakistan, whereas the data for the 2025 events covers large regions of Pakistan. A table of these events and the periods to which they refer to are displayed below:

Event/Data Period Covered Baseline Date
2022 Floods (covers majority of areas administered by Pakistan) 14 August – 7 September 2022 30 June 2022
2025 Floods - Punjab (covers areas in and around Punjab province) 19 August - 28 August 2025 5 July 2025
2025 Floods - Sindh (covers areas in and around Sindh province) 21 August - 1 Sepetember 2025 7 July 2025
  • The data for each event consist of two datasets, Population During Crisis and Movement During Crisis:

    • Population – these are population stock data, showing the number of users in each spatial unit at three snapshots: 00:00, 8:00 and 16:00 (Pacific Time). The data are removed when there are fewer than 10 observations.

    • Movement – these are population flow data, showing the origin and destination of users between temporal points. Users’ origin and destination are chosen according to where they spent most time within each 8-hour window. For example, data recorded at 16:00 shows the flow between areas from 08:00 to 16:00. Where there are fewer than 10 observations for a flow, data are removed. The Pakistan data available for the Movement data are only at the 08:00 and 16:00 time periods, but not for the 00:00 period.

  • The data are at two geographic scales: 800m tiles (known as quadkeys, based on the Bing Maps tile system) and aggregated to global administrative (GADM) level 2 geographies.

  • The data are generally comprehensive, but there are some gaps. For example, as stated above, in both the tiled and aggregated Movement data, only the 08:00 and 16:00 time stamps are available, with the 00:00 period missing. Additionally, the aggregated Population data for 20, 21 and 22 August 2022 are missing, and these data are not available at all for the 2025 flood event in and around Karachi, though they are available for the 800m tiles.

  • Both datasets contain data from a baseline period before the disaster event to compare users’ stock or flow during the crisis. The raw and percentage differences are provided within the dataset, along with a z-score to assess the statistical significance of the change from the pre-crisis baseline to the crisis period.

  • The data is available as a series of csv files. Each file is either the tiled or aggregated data for each time stamp of each day of the Population or Movement data.

A guide on using these data in R can be found here.

2.2 IOM Data

2.3 Contextual Data

  • Alongside mobility data, we have prepared a series of datasets that provide additional context to Pakistan. They consist of population data in raster and aggregated form, and socioeconomic data, also in raster and aggregated format.

  • Raster population data:

    • Population estimates for Pakistan for 2020 from WorldPop. The data is at 100m resolution grids in a single raster file.

    • Population estimates for Pakistan by age and sex for 2020 from WorldPop. The data is at 100m resolution grids in a series of rasters. Files are structured like {iso} {gender} {age group} {year} {type} {resolution}.tif - gender fields are f (female), m (male), t (total); age group fields are 00 (0-12 months), 01 (1-4 years), 05 (5-9 years) and so on until 90 (age 90 and above).

  • Aggregated population data:

    • Population estimates for Pakistan for 2020 from the WorldPop raster, aggregated (total_pop) to global administrative level 2, and the province and district polygons used by OCHA. The data is calculated by the project team using WorldPop data and spatial polygons provided by GADM and OCHA.
  • Socioeconomic data:

    • Deprivation data from The Global Gridded Relative Deprivation Index (GRDI). The data is a raster file at 1km resolution, cropped to Pakistan. Data are on a 0-100 scale, with high values indicating higher relative levels of deprivation. Complete documentation for the data can be accessed here. These have also been aggregated to global administrative level 2 and the province and district polygons used by OCHA by computing the mean value (mean_rdi) within the polygon from the GRDI raster data.
  • Flood data:

    • Satellite detected water extents in Pakistan between 01 and 29 August 2022. The data is a series of shapefiles showing flood extent based on satellite imagery. Data is from the UN Operational Satellite Applications Programme (UNOSAT), found here.

    • Satellite detected water extents in Pakistan between 26 August and 07 September 2025. The data is a series of shapefiles showing flood extent based on satellite imagery. Data is from the UN Operational Satellite Applications Programme (UNOSAT), found here.

  • Rainfall data:

    • Raster accumulated precipitation data from the NASA Global Precipitation Measurement for Pakistan. Data is daily accumulated precipitation in mm at 10km resolution at two time periods: 01 July - 07 September 2022 and 01 July - 07 September 2025. Data is specifically the late run dataset, available in near-real-time with a ~12 hour delay. 
  • Climate data:

    • Dekadal (10-day) rainfall data from 2022 onward at the district level (coded to OCHA administrative level 2) sourced from the World Food Programme (WFP). Data shows series of metrics for rainfall, listed here.

    • Dekadal (10-day) Normalized Difference Vegetation Index (NDVI) data from 2022 onward at the district level (coded to OCHA administrative level 2) sourced from the World Food Programme (WFP). Data shows series of metrics for vegetation greenness, listed here, normally used to quantify the health and density of vegetation.

2.4 Spatial Boundaries

  • Due to recent boundary changes and differing administrative boundaries being used by different organisations, the joining of spatial data together for Pakistan is not a simple task. Much of the data listed above, however, is able to be joined to a spatial boundary to be mapped and analysed. This section describes the spatial boundaries made available, as well a table with details of how the Meta, IOM and contextual datasets can be joined to these.

  • The boundaries available are:

    • Spatial polygons for global administrative levels, sourced from GADM. The shapefile boundaries for global administrative levels 0, 1, 2 and 3 and the geopackage for Pakistan are included. These boundaries can be joined to the aggregated Facebook Population and Movement data, as well as the population and deprivation data aggregated to global administrative level 2.

    • Spatial polygons used by OCHA for subnational boundaries of Pakistan. Data is shapefiles for provinces (adm1), districts (adm2) and tehsils (adm3), sourced from OCHA. These boundaries can be joined to the spatial variable codes in the IOM CNI data for provinces and districts, as well as the population and deprivation data aggregated to OCHA administrative level 1 and 2, and to the WFP climate data.

    • Spatial polygons for the 800m tiles (also known as quadkeys) found in the Facebook Population and Movement tile data. These were created from the tile datasets using the R package quadkeyr.

  • The below table provides an overview of the variables in each dataset and which boundary they relate to:

Dataset Spatial Codes Type Boundary File and Corresponding Code
Facebook Data
Facebook Population (aggregated) polygon_id GADM level 2 Polygon Global administrative level 2 polygons (GID_2)
Facebook Population (tiled) quadkey 800m tile Quadkey polygons (quadkey)
Facebook Movement (aggregated)

start_polygon_id

end_polygon_id

Origin polygon

Destination polygon

Global administrative level 2 polygons (GID_2)
Facebook Movement (tiled)

start_quadkey

end_quadkey

Origin 800m tile

Destination 880m tile

Quadkey polygons (quadkey)
IOM Data
IOM CNI Data

ProvinceCode/Pcode/PCODE; or Admin1 pcode

District Code/Pcode/PCODE; or Admin2 pcode

Province (adm1) polygon

District (adm2) polygon

OCHA subnational polygons (ADM1_PCODE) (ADM2_PCODE)
Contextual Data
Aggregated Population Data (gadm_2) GID_2 GADM level 2 Polygon Global administrative level 2 polygons (GID_2)
Aggregated Population Data (ocha_1 and ocha_2)

ADM1_PCODE

ADM2_PCODE

Province (adm1) polygon

District (adm2) polygon

OCHA subnational polygons (ADM1_PCODE)(ADM2_PCODE)
Aggregated Global Relative Deprivation Index Data (GRDI) (gadm_2) GID_2 GADM level 2 Polygon Global administrative level 2 polygons (GID_2)
Aggregated Global Relative Deprivation Index Data (GRDI) (ocha_1 and ocha_2)

ADM1_PCODE

ADM2_PCODE

Province (adm1) polygon

District (adm2) polygon

OCHA subnational polygons (ADM1_PCODE)(ADM2_PCODE)
WFP - Rainfall at Subnational Level ADM2_PCODE District (adm2) polygon OCHA subnational polygons (ADM2_PCODE)
WFP - NVDI at Subnational Level ADM2_PCODE District (adm2) polygon OCHA subnational polygons (ADM2_PCODE)

Other Resources

Although not the main focus of the FloodTraces Hackathon, participants may find it helpful to check out work doene as part of the Empowering Resilience in a Sinking City project, which focused on the case of Jakarta, Indonesia.

You can learn more about the project Sinking City and access the data they used in their own Hackathon event here.