-
Notifications
You must be signed in to change notification settings - Fork 0
Rough draft of satellite / active fire #4
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
benartuso
wants to merge
4
commits into
main
Choose a base branch
from
add-active-fire-satellite
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
4 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,172 @@ | ||
|
|
||
| # Preparing Satellite Data for Wildfire Detection with Oxen.ai 🔥🛰️ | ||
|
|
||
| ## Introduction | ||
| Satellite imagery data is crucial to solving many of the worlds problems today. Wildfires is one such problem in which we can unleash the power of machine learning to help solve. Extreme wildfires can result in the loss of lives, ecosystems and entire communities. This costs societies billions of dollars to rebuild, and the earth years to recover from. | ||
|
|
||
| With the help of machine learning and computer vision, we can help improve the detection of wildfires, enabling faster response times. Autonomous vehicles can also assist in suppressing high risk fires in dangerous places for humans to navigate. | ||
|
|
||
| Luckily, there are few domains with greater quantity of high-quality, publicly accessible, data than satellite imagery. Government-funded missions such as [LANDSAT](https://www.usgs.gov/landsat-missions#:~:text=Since%201972%2C%20Landsat%20satellites%20have,natural%20resources%20and%20the%20environment.) (USGS / NASA) and [SENTINEL](https://sentinel.esa.int/web/sentinel/missions/sentinel-2) (ESA) have captured and transmitted terabytes of this data every day for multiple decades, providing an extremely rich resource to those wishing to develop models for pressing geospatial challenges like disaster response. | ||
|
|
||
|  | ||
|
|
||
| While these datasets are information-rich and freely available, they are notoriously difficult to collaborate on and work with. | ||
|
|
||
| ## Making The Data More Accessible | ||
| In this tuturial, we will be working with a dataset that has meticulously annotated by researchers for space-based active wildfire detection. We'll be walking through how to prepare the [Landsat-8 Imagery](https://drive.google.com/drive/folders/1GIcAev09Ye4hXsSu0Cjo5a6BfL3DpSBm) (introduced in [this paper by Pereira et. al](https://arxiv.org/abs/2101.03409)) | ||
|
|
||
| The data was originally stored in Google Drive, making it hard to explore and collaborate on, since it is a collection of zip files that need to be downloaded and processed. | ||
|
|
||
|  | ||
|
|
||
| In this tutorial, we'll use tooling from [Oxen.ai](www.oxen.ai) to address the common pitfalls in working with satellite imagery. In doing so, we'll transform this widely-cited dataset for wildfires detection from a 10GB Google Drive-hosted .zip file collection to a cloneable, well-organized [Oxen.ai](www.oxen.ai) data repository ready for collaborative model development. | ||
|
|
||
| TODO: Add a beautiful image of the oxen UI | ||
|
|
||
| ## The LANDSAT-8 Satellite Imagery Dataset | ||
|
|
||
| This dataset contains patches of imagery extracted from LANDSAT-8 satellite imagery across Earth, many of which contain actively burning wildfires. | ||
|
|
||
| While the authors processed the full 1.6TB of imagery captured by LANDSAT-8 in August 2020 (targeting peak fire season), we're primarily interested in the excellent subset of data for which they manually created segmentation masks denoting active burn. | ||
|
|
||
|  | ||
|
|
||
| In addition to being a more manageable (12GB) dataset size for model fine-tuning and evaluation applications, these human-verified labels will give any downstream models the best possible signal with which to learn to detect fires in novel imagery. | ||
|
|
||
|  | ||
|
|
||
|
|
||
| To supplement these ground-truth manual annotations, the researchers include 5 auto-generated segmentation masks that don't require any direct human input. These are generated computationally, using algorithms previously developed and validated in the remote sensing literature. While less precise than the manual annotations, these can also be a useful reference point for model validation. | ||
|
|
||
| The authors originally use a [GitHub repository](https://github.com/pereira-gha/activefire/tree/main) to distribute the dataset, but the only data currently hosted there are a 77-row metadata file called `images2020009.csv` (we'll come back to this shortly) and some download scripts to fetch the full dataset from .zip files Google Drive. | ||
|
|
||
|  | ||
|
|
||
|
|
||
| ### 🐂 Oxen is better than zip files | ||
| We see two primary areas in which we can better ready this dataset for collaborative model development. | ||
|
|
||
| **1. Enable atomic changes.** If we want to add, modify, or remove 50 MB of data from this dataset, we should only have to push (and our colleagues subsequently pull) 50 MB — not the whole 10 GB zip file. | ||
|
|
||
| If we identify a mistake and want to roll these changes back, we shouldn't have to hunt for the `DATA_VERSION_3_FINAL.zip` taking up extra storage on our drive - instead, we'll leverage Oxen's version control to seamlessly undo the change. | ||
|
|
||
| **2. Improve data-label linking.** The source data and 6 human + computed label sets are currently linked somewhat tenuously through pattern matching in the image file paths. We'll nail this down to an authoritative, well-validated data -> labels mapping file to save our teammates some meticulous ETL work and speed up the model development process. | ||
|
|
||
|
|
||
| # TODO: Make the following more step-by-step-y | ||
|
|
||
| Below feels pretty bloated and repetitive with what we talked about earlier. Can we just walk them through step by step how to transfer, with lightly touching on the benefits of each step? I imagine one or two sentences per step, and one image per step to keep the pace up and scannability high. | ||
|
|
||
| ### Making the data modeling-ready with Oxen | ||
|
|
||
| #### 1. **Enabling atomic, versioned dataset collaboration** | ||
| We want to enable our team to efficiently iterate on this data throughout the duration of the wildfire detection project. In this dataset's original distribution via a collection of .zip files hosted on Google Drive, any changes, however small, to the source data would require a recompression and re-upload of the 10GB data folder. | ||
|
|
||
|  | ||
|
|
||
| Integration of these changes by other team members would then require a re-download of the full dataset and manual merging with their local work, without any built-in awareness of the changes made between versions. | ||
|
|
||
| Instead, we can unpack this data into an Oxen repository so that all our teammates can easily: | ||
| - Examine, query and explore the data in the OxenHub UI | ||
| - Clone it locally with `oxen clone https://www.hub.oxen.ai/ba/ActiveFire` | ||
| - Create a space-saving local working copy with `oxen checkout -b my-local-branch` | ||
| - Contribute to the improvement of the repo through adding, deleting, and modifying additional files without needing to re-package and re-push all 10GB of data. | ||
| - TODO: remote staging? or too much? | ||
|
|
||
| **Early exploration and understanding** | ||
|
|
||
| While waiting for the .zip files to download and unpack from the Google Drive source, we created a README referencing back to the original project. | ||
|
|
||
| Included with the data was a tabular file titled `images202009.csv`—we couldn't quite tell what this file represented at first, so we pushed it up to Oxen as well to explore further. | ||
|
|
||
|  | ||
| Already, we're getting a lot more information about the structures and datatypes of the dataset, while adding a version-controlled history we can turn to if we make any mistakes along the way. | ||
|
|
||
| Diving a bit further into the images202009 file... | ||
|  | ||
| ...OxenHub gives us no-fuss insight into the shape and schema of the dataset, and the ability to explore it further to understand what it represents. | ||
|
|
||
| The `productId` column ("LC08") and `cloudCover` tipped us off that this is a metadata export of full 5600x6100 pixel LANDSAT scenes from which the researchers exported their 256x256 pixel patches for modeling. | ||
|
|
||
|  | ||
|
|
||
|  | ||
|
|
||
| We can use the natural language query interface to translate our analytic questions into SQL to reach a better understanding of which terrain correction methods are present in the dataset. The imbalance here between Precision and Terrain Correction (L1TP) and Systematic Terrain Correction (L1GT) is important to flag for our team during model and data development, as they could cause inconsistencies in the resulting data or confusion in the model training process. | ||
|
|
||
| **Unpacking and restructuring** | ||
| Now that the Google Drive downloads have unzipped, we can `oxen add`, `oxen commit`, and then `oxen push` them to our repository, first in their original format... | ||
|  | ||
| ...then, after a bit of friction, in a reorganized format that will set us up for clearer delineation and linkage between our data and labels: | ||
|
|
||
|  | ||
|
|
||
|
|
||
| #### 2. Improving data-label linking and training ergonomics | ||
| In the source dataset, the original satellite image for a patch, the authors' manually annotated labels, and the 5 sets of algorithmically determined labels are linked together only implicitly through commonalities in the filenames: | ||
|
|
||
| ``` | ||
| Input data (LANDSAT) path: | ||
| - LC08_L1GT_226074_20200921_20200921_01_RT_p00811 | ||
|
|
||
| Manual annotation path: | ||
| - LC08_L1GT_226074_20200921_20200921_01_RT_v1_p00811 | ||
|
|
||
| Kumar-Roy: | ||
| - LC08_L1GT_226074_20200921_20200921_01_RT_Kumar-Roy_p00811 | ||
|
|
||
| Murphy: | ||
| - LC08_L1GT_226074_20200921_20200921_01_RT_Murphy_p00811 | ||
|
|
||
| Intersection: | ||
| - LC08_L1GT_226074_20200921_20200921_01_RT_Intersection_p00811 | ||
|
|
||
|
|
||
| ``` | ||
|
|
||
| While this isn't particularly difficult to parse with regular expressions at training or inference time, it's also an opportune hiding spot for sneaky human error and unnecessary data munging overhead on each of the individual engineers working on this project. | ||
|
|
||
| Further, while there are 9,044 total patches of raw imagery included in this dataset, the various annotation types vary in number and don't cover the same data: | ||
| - Manual annotations: 100 | ||
| - Kumar method: 391 | ||
| - Murphy method: 164 | ||
| - Schroeder method: 227 | ||
| - Intersection method: 118 | ||
| - Voting method: 198 | ||
|
|
||
| ...so solidifying awareness of coverage across the various labeling types in one place is key to effective collaboration on this dataset. | ||
|
|
||
| **Creating a data -> label mapping file and validating with OxenHub** | ||
|
|
||
| We parsed the filenames to the common patch_id, creating the following mapping file structure... | ||
|
|
||
|
|
||
|  | ||
| ...associating each LANDSAT patch to the various label paths for that patch, when they exist. | ||
|
|
||
| We can use schema metadata to let OxenHub know which columns are relative filepaths: | ||
|
|
||
| ``` | ||
| TODO: Schema setting from CLI | ||
| ``` | ||
|
|
||
| which quickly elucidates an error in our processing. | ||
|
|
||
| **TODO IMAGE OF FAILED FILE EXISTENCE CHECK IN HUB** | ||
|
|
||
| After hunting down the bug in our filepath processing code, we make a new commit with the fix and are ready to train! | ||
|
|
||
| **TODO IMAGE OF SUCCESSFUL FILE EXISTENCE CHECKS IN HUB** | ||
|
|
||
| This single source of truth for relating files across our repository will enable our team to more confidently investigate, sample, and evaluate on this dataset without having to worry about the integrity of the data-label linkage. | ||
|
|
||
| ### And we're off! | ||
| With these reorganizations, validations, and improvements, this dataset is now much more readily usable for collaborative model development for wildfire detection and segmentation applications. | ||
|
|
||
| We'd love to see what kind of models you're building on top of it! Clone the repo [here](https://www.oxen.ai/ba/ActiveFire), reach out at [hello@oxen.ai](mailto:hello@oxen.ai), follow us on Twitter [@oxendrove](https://twitter.com/oxendrove?ref=blog.oxen.ai), dive deeper into the [documentation](https://github.com/Oxen-AI/oxen-release?ref=blog.oxen.ai), or **Sign up for Oxen today. [http://oxen.ai/register.](http://oxen.ai/register.?ref=blog.oxen.ai)** | ||
|
|
||
| And remember—for every star on [GitHub](https://github.com/Oxen-AI/oxen-release?ref=blog.oxen.ai), an ox gets its wings. | ||
|
|
||
| No, really...we hooked up an [Oxen repo](https://oxen.ai/ox/FlyingOxen?ref=blog.oxen.ai) to a GitHub web-hook that runs Stable Diffusion every time we get a star. [Go find yours!](https://oxen.ai/ox/FlyingOxen?ref=blog.oxen.ai) | ||
|
|
||
|  | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Can we put these images side by side? Feels like a lot of scrolling.
Also look at you being so nice to future generative AI data crawls with the alt text 😉