Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file added Nguyen & Cao_Final Project Upds.pptx
Binary file not shown.
9,339 changes: 9,339 additions & 0 deletions Nguyen + Cao Final Project/Data400_FinalProject.ipynb

Large diffs are not rendered by default.

Binary file added Nguyen + Cao Final Project/Giant.xlsm
Binary file not shown.
Binary file not shown.
Binary file added Nguyen + Cao Final Project/Target.xlsm
Binary file not shown.
Binary file added Nguyen + Cao Final Project/Walmart.xlsx
Binary file not shown.
64 changes: 64 additions & 0 deletions Nguyen + Cao Final Project/readme.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
├── Data
│ ├── Walmart.xlsx
│ ├── Target.xlsx
│ └── Giant.xlsx
├── Scripts
│ └── Analysis.ipynb
├── Outputs
│ ├── Visualizations
│ ├── Multi-linear regression model
│ └── Hyperparameter tuning
└── README.md

Same Product, Different Price: A Comparative Study of Retail Pricing Strategies
I. Overview:
- This project investigates how the same branded groceries products are priced differently across major U.S. retailers, specifically Walmart, Target, and Giant. By analyzing pricing patterns, the study aims to
uncover whether certain stores consistently offer lower/higher prices, and if pricing varies by product category or unit size. These findings may inform consumer behavior insights and shed light on each store’s
pricing strategy.

II. Directory Structure
1. Data
a. Walmart.xlsx
b. Target.xlsx
c. Giant.xlsx

2. Scripts
a. Analysis.ipynb

3. Outputs
a. Visualizations
b. Multi-linear regression model
c. Hyperparameter tuning

4. README.md

III. Data Sources
- Produced listings and prices exported from each retailer’s website in Spring 2025
- Key columns across Walmart, Target, Giant datasets:
o Product_url_link
o Product_brand
o Product_type
o Product_price_per_unit
o Product_price

IV. Main Processing Steps
1. Load and clean Products Data
• For each retailer (Walmart, Target, Giant), data from the groceries categories is loaded and cleaned in Excel
• Standardize column names, units, and price formats; remove null or product brands irrelevant to the initial list
2. Normalize unit prices
• Convert all pricing to standard per-ounce in dollar format for fair comparison
3. Fuzzy match products
• Install thefuzz package and use thefuzz to identify samebranded, same quantity across retailers
4. Statistical analysis
• EDA
• Linear regression
• Hyperparameter tuning
• Visualizations

V. Installation/Requirements
- Excel
- Python language
- Pandas, numpy, matplotlib, thefuzz, seaborn, sklearn.linear_model

VI. References
- Cite sources: Walmart.com, Target.com, GiantFood.com
Binary file added Presentation 1.pdf
Binary file not shown.
15 changes: 15 additions & 0 deletions Project Description.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Project Description
My mini-project will explore the environmental impacts on rat presence in New York City. I plan to use rat inspection data as the primary inputs of the model and join environmental data like temperature, weather, seasons, etc to inform the model. This outputs will include preddictions of rat sightings based on these factors and ultimately be used to generate predictions and examine hotspots in the city. This will allow city officials and pest companies to target thir surveillance and thus treatment to reduce this issue in NYC.

# Dataset
The main dataset of this project will be the Rodent Inspection Dataset from NYC Open Data: https://data.cityofnewyork.us/Health/Rodent-Inspection/p937-wjvj/data. I will also be web scraping weather/temperature/seasons data to support these pursuits. Lastly, I will also find the shapefile version of the rodent insepction data and join the outputs to create a hotspots/targetted map.

# Methodologies
a. Data Cleaning: I plan to go through and subset the data to only include places with rodent activity/presence. Clean and compile weather/environmental data from different sources pertaining to different negihborhoods in NYC. Scrape any data relevant to funding sources for the rodent issue.
b. Data Wrangling: Join environmental monitoring data with Rodent Insepction Datasheet. Clean and filter accordingly
c. Modeling: Split and train dataset with a Regression or time-series model.
d. Tune: Tune hyperparameters and adjust features based on the results.
e. Combine and import to ArcGIS for hotspot mapping.

# Impact on Stakeholders
This project will impact communities in NYC and the policymakers who are working to address this issue. Rodent infestations in the city is a big hygiene and thus public health issue. This project will identify hotspots of rodents in the city and find ways to stretch the budget and allocate funding and resources to combat the areas impacted most.
Binary file added cao m - 400 mini proj.pptx
Binary file not shown.
28 changes: 28 additions & 0 deletions mcao_idea2.md

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great idea to help the farm expand its distribution to students as well as potentially expanding into other markets. It’s great that the college farm keeps track of many data categories. A linear regression model may be useful in predicting how significant certain categories impact sales and overall inventory values.

Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
Project Title: Optimize Meals and Profits at Dickinson Farmworks

Below is the proposal for John Park and I's DATA 300 project. I am submitting this as a backup project plan with the following revisions:

More parameters (temp, season, break down soups sold into 15 min intervals)
One hot encoding instead of label encoding
Perform tree and regression models again
Visualize the data using PowerBI/Tableau

Project Description

Farmworks is a farm-to-table store that prepares meals from fresh ingredients supplied by the Dickinson College Farm. This business serves Dickinson students and faculty and the greater Carlisle area to aim zero-waste emission within the store and strengthen the community. A recent conversation with Jenn Halpin, Director of Dickinson College Farm, suggested that Farmworks is reaching capacity in the number of meals they can produce each day and are therefore running out of food before the business closes. Our goal is to optimize the number of soups and salads the head chef should make, in order to maximize their profits (of each business day 11a to 2p). To pursue this project, we will collect historical data on menu items, amount of food produced and served, cost to produce each meal, timestamp, and potentially weather data to inform the model. Additionally, we will consider the number of workers to scale its influence on the production side.

We plan to apply linear regression and decision trees models to identify significant various factors and generate how many more meals need to be produced. The model will find a balance between the needs of the community and maximize the business profits. We are hopeful that the outputs of the project will inform Farmwork's’ discussion with college leadership to expand budget for another chef and attract more customers by consistently satisfying the demand.


Dataset

We are in the process of collecting data from the Farmworks’ business system (Square) that contains time, date, # of meals served, what they bought (salad, soup, ½ soup, ½ salad or empanadas). We will also receive data on the base estimate cost of each soup and information related to menu rotations and food preparation. We may also choose to include weather data and see how that feature informs our model.

For specific features, we plan to include the historical menu data (categorical), number of meals sold (numerical, continuous), special menu items (e.g. empanadas, breakfast, etc.) (categorical), day of the week (categorical), number of servings made (categorical), # of servings sold per hour and estimate cost (numerical), and potentially number of workers/total labor hours (numerical) and weather data (numerical (temp), categorical (cloudy, raining, etc.)).


Evaluation

To evaluate our model’s performance in predicting optimal meal production, we will use regression metrics, including Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R-squared values (R^2). The MAE and RMSE metrics will assist in identifying and adjusting significant prediction gaps. R-squared will determine our model’s effectiveness in capturing key demand features. If it leads to overfitting, we may switch to the adjusted R-squared metric. These metrics will assess the model’s accuracy in predicting the meal quantities, classified by each menu cycle, and effectiveness in meeting the equilibrium point of the meal’s input cost and revenue. After the initial evaluation, we will tune hyperparameters and refine features, especially weather and labor productivity, to improve the model’s performance. The model’s sensitivity to weather and seasonal menu changes will be ensured to maintain a consistent performance.

After collecting and performing initial tests, we may potentially increase the features and experiment with decision trees.
64 changes: 64 additions & 0 deletions readme
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
├── Data
│ ├── Walmart.xlsx
│ ├── Target.xlsx
│ └── Giant.xlsx
├── Scripts
│ └── Analysis.ipynb
├── Outputs
│ ├── Visualizations
│ ├── Multi-linear regression model
│ └── Hyperparameter tuning
└── README.md

Same Product, Different Price: A Comparative Study of Retail Pricing Strategies
I. Overview:
- This project investigates how the same branded groceries products are priced differently across major U.S. retailers, specifically Walmart, Target, and Giant. By analyzing pricing patterns, the study aims to
uncover whether certain stores consistently offer lower/higher prices, and if pricing varies by product category or unit size. These findings may inform consumer behavior insights and shed light on each store’s
pricing strategy.

II. Directory Structure
1. Data
a. Walmart.xlsx
b. Target.xlsx
c. Giant.xlsx

2. Scripts
a. Analysis.ipynb

3. Outputs
a. Visualizations
b. Multi-linear regression model
c. Hyperparameter tuning

4. README.md

III. Data Sources
- Produced listings and prices exported from each retailer’s website in Spring 2025
- Key columns across Walmart, Target, Giant datasets:
o Product_url_link
o Product_brand
o Product_type
o Product_price_per_unit
o Product_price

IV. Main Processing Steps
1. Load and clean Products Data
• For each retailer (Walmart, Target, Giant), data from the groceries categories is loaded and cleaned in Excel
• Standardize column names, units, and price formats; remove null or product brands irrelevant to the initial list
2. Normalize unit prices
• Convert all pricing to standard per-ounce in dollar format for fair comparison
3. Fuzzy match products
• Install thefuzz package and use thefuzz to identify samebranded, same quantity across retailers
4. Statistical analysis
• EDA
• Linear regression
• Hyperparameter tuning
• Visualizations

V. Installation/Requirements
- Excel
- Python language
- Pandas, numpy, matplotlib, thefuzz, seaborn, sklearn.linear_model

VI. References
- Cite sources: Walmart.com, Target.com, GiantFood.com