Human body segmentation for CS 191 (Senior Research Project).
| Directory | Description |
|---|---|
sample-backpack/ |
Sample images from MS COCO |
scripts/ |
Scripts for testing, data processing, regularization, etc. |
assets/ |
Poster, images, and resources describing this project |
assets/output_imgs/ |
Predicted segmentation masks for COCO samples in sample-backpack/ |
model/ |
Diamondback model architecture, losses, DenseNet encoder |
logs/ |
Keras training logs over 19 epochs (zipped) |
util/ |
Data generators, common paths |
weights/ |
Location where Diamondback model weights are stored and loaded from |
tmp/ |
Misc directory for debugging, used in test scripts |
Download the model weights (not in GitHub because of size):
python3 download_diamondback_weights.py
# -> weights/diamondback_{...}.h5
# -> model/densenet_encoder/encoder_model.h5Run a demo prediction:
python3 predict.py --demo --load_path weights/diamondback_{...}.h5Train:
python3 training.py [--debug] [--load_path weights/diamondback_{...}.h5]Note
To run certain data scripts, make sure to install (or be in a virtualenv with) the COCO API.
Some paths (namely util/pathutil.py, shell scripts in tmp/) are hardcoded for FloydHub's environment, where data and model weights were expected to be at the drive root /.
Inspiration mostly from Fu et al. in SDN for Semantic Segmentation (2017). We trained Diamondback M2, which is two encoder-decoder units. More units add parameters but may improve results.
- Each unit employs dense convolutional connections.
- Inter-unit connections: Decoder feature maps at unit N-1 are concat'd with the encoder feature maps of the same resolution at unit N.
- DenseNet encoder: We transfer learn by using DenseNet-161's layers as the first unit's encoder.
- 28x28 and 56x56 feature maps from the DN encoder are convolved and concat'd to every other unit's decoder layers of the respective resolution.
- Using all units' learned decoder outputs: We concatenate all units' decoder outputs (56x56), then upsample up to 224x224 for a final prediction tensor with 2 channels.
Losses and IOU over 19 epochs of training.
| Input | Prediction | Ground Truth |
|---|---|---|
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |


















































