-
Notifications
You must be signed in to change notification settings - Fork 32
mini project idea #35
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
mcao694
wants to merge
13
commits into
ernbilen:main
Choose a base branch
from
mcao694:main
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
13 commits
Select commit
Hold shift + click to select a range
7cc666d
mini project idea
mcao694 b37b088
idea 2 backup plan for mini project
mcao694 2c73d67
Presentation 200
mcao694 1f90e06
mini proj slide deck
mcao694 d7fee5b
final project update
mcao694 9aa1adb
more data
mcao694 0ecc507
Delete shit.ipynb
mcao694 e7d0672
Delete walmart1.xlsx
mcao694 cb7180c
Delete Giant.xlsm
mcao694 dc6e63f
Delete Target.xlsm
mcao694 4dd2610
Create readme
mcao694 8643d0b
Update readme
mcao694 34a2501
final project data + presentation + readme
mcao694 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
Binary file not shown.
9,339 changes: 9,339 additions & 0 deletions
9,339
Nguyen + Cao Final Project/Data400_FinalProject.ipynb
Large diffs are not rendered by default.
Oops, something went wrong.
Binary file not shown.
Binary file added
BIN
+1.43 MB
Nguyen + Cao Final Project/Nguyen & Cao - DATA 400 Final Project.pptx
Binary file not shown.
Binary file not shown.
Binary file not shown.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,64 @@ | ||
| ├── Data | ||
| │ ├── Walmart.xlsx | ||
| │ ├── Target.xlsx | ||
| │ └── Giant.xlsx | ||
| ├── Scripts | ||
| │ └── Analysis.ipynb | ||
| ├── Outputs | ||
| │ ├── Visualizations | ||
| │ ├── Multi-linear regression model | ||
| │ └── Hyperparameter tuning | ||
| └── README.md | ||
|
|
||
| Same Product, Different Price: A Comparative Study of Retail Pricing Strategies | ||
| I. Overview: | ||
| - This project investigates how the same branded groceries products are priced differently across major U.S. retailers, specifically Walmart, Target, and Giant. By analyzing pricing patterns, the study aims to | ||
| uncover whether certain stores consistently offer lower/higher prices, and if pricing varies by product category or unit size. These findings may inform consumer behavior insights and shed light on each store’s | ||
| pricing strategy. | ||
|
|
||
| II. Directory Structure | ||
| 1. Data | ||
| a. Walmart.xlsx | ||
| b. Target.xlsx | ||
| c. Giant.xlsx | ||
|
|
||
| 2. Scripts | ||
| a. Analysis.ipynb | ||
|
|
||
| 3. Outputs | ||
| a. Visualizations | ||
| b. Multi-linear regression model | ||
| c. Hyperparameter tuning | ||
|
|
||
| 4. README.md | ||
|
|
||
| III. Data Sources | ||
| - Produced listings and prices exported from each retailer’s website in Spring 2025 | ||
| - Key columns across Walmart, Target, Giant datasets: | ||
| o Product_url_link | ||
| o Product_brand | ||
| o Product_type | ||
| o Product_price_per_unit | ||
| o Product_price | ||
|
|
||
| IV. Main Processing Steps | ||
| 1. Load and clean Products Data | ||
| • For each retailer (Walmart, Target, Giant), data from the groceries categories is loaded and cleaned in Excel | ||
| • Standardize column names, units, and price formats; remove null or product brands irrelevant to the initial list | ||
| 2. Normalize unit prices | ||
| • Convert all pricing to standard per-ounce in dollar format for fair comparison | ||
| 3. Fuzzy match products | ||
| • Install thefuzz package and use thefuzz to identify samebranded, same quantity across retailers | ||
| 4. Statistical analysis | ||
| • EDA | ||
| • Linear regression | ||
| • Hyperparameter tuning | ||
| • Visualizations | ||
|
|
||
| V. Installation/Requirements | ||
| - Excel | ||
| - Python language | ||
| - Pandas, numpy, matplotlib, thefuzz, seaborn, sklearn.linear_model | ||
|
|
||
| VI. References | ||
| - Cite sources: Walmart.com, Target.com, GiantFood.com |
Binary file not shown.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,15 @@ | ||
| # Project Description | ||
| My mini-project will explore the environmental impacts on rat presence in New York City. I plan to use rat inspection data as the primary inputs of the model and join environmental data like temperature, weather, seasons, etc to inform the model. This outputs will include preddictions of rat sightings based on these factors and ultimately be used to generate predictions and examine hotspots in the city. This will allow city officials and pest companies to target thir surveillance and thus treatment to reduce this issue in NYC. | ||
|
|
||
| # Dataset | ||
| The main dataset of this project will be the Rodent Inspection Dataset from NYC Open Data: https://data.cityofnewyork.us/Health/Rodent-Inspection/p937-wjvj/data. I will also be web scraping weather/temperature/seasons data to support these pursuits. Lastly, I will also find the shapefile version of the rodent insepction data and join the outputs to create a hotspots/targetted map. | ||
|
|
||
| # Methodologies | ||
| a. Data Cleaning: I plan to go through and subset the data to only include places with rodent activity/presence. Clean and compile weather/environmental data from different sources pertaining to different negihborhoods in NYC. Scrape any data relevant to funding sources for the rodent issue. | ||
| b. Data Wrangling: Join environmental monitoring data with Rodent Insepction Datasheet. Clean and filter accordingly | ||
| c. Modeling: Split and train dataset with a Regression or time-series model. | ||
| d. Tune: Tune hyperparameters and adjust features based on the results. | ||
| e. Combine and import to ArcGIS for hotspot mapping. | ||
|
|
||
| # Impact on Stakeholders | ||
| This project will impact communities in NYC and the policymakers who are working to address this issue. Rodent infestations in the city is a big hygiene and thus public health issue. This project will identify hotspots of rodents in the city and find ways to stretch the budget and allocate funding and resources to combat the areas impacted most. |
Binary file not shown.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,28 @@ | ||
| Project Title: Optimize Meals and Profits at Dickinson Farmworks | ||
|
|
||
| Below is the proposal for John Park and I's DATA 300 project. I am submitting this as a backup project plan with the following revisions: | ||
|
|
||
| More parameters (temp, season, break down soups sold into 15 min intervals) | ||
| One hot encoding instead of label encoding | ||
| Perform tree and regression models again | ||
| Visualize the data using PowerBI/Tableau | ||
|
|
||
| Project Description | ||
|
|
||
| Farmworks is a farm-to-table store that prepares meals from fresh ingredients supplied by the Dickinson College Farm. This business serves Dickinson students and faculty and the greater Carlisle area to aim zero-waste emission within the store and strengthen the community. A recent conversation with Jenn Halpin, Director of Dickinson College Farm, suggested that Farmworks is reaching capacity in the number of meals they can produce each day and are therefore running out of food before the business closes. Our goal is to optimize the number of soups and salads the head chef should make, in order to maximize their profits (of each business day 11a to 2p). To pursue this project, we will collect historical data on menu items, amount of food produced and served, cost to produce each meal, timestamp, and potentially weather data to inform the model. Additionally, we will consider the number of workers to scale its influence on the production side. | ||
|
|
||
| We plan to apply linear regression and decision trees models to identify significant various factors and generate how many more meals need to be produced. The model will find a balance between the needs of the community and maximize the business profits. We are hopeful that the outputs of the project will inform Farmwork's’ discussion with college leadership to expand budget for another chef and attract more customers by consistently satisfying the demand. | ||
|
|
||
|
|
||
| Dataset | ||
|
|
||
| We are in the process of collecting data from the Farmworks’ business system (Square) that contains time, date, # of meals served, what they bought (salad, soup, ½ soup, ½ salad or empanadas). We will also receive data on the base estimate cost of each soup and information related to menu rotations and food preparation. We may also choose to include weather data and see how that feature informs our model. | ||
|
|
||
| For specific features, we plan to include the historical menu data (categorical), number of meals sold (numerical, continuous), special menu items (e.g. empanadas, breakfast, etc.) (categorical), day of the week (categorical), number of servings made (categorical), # of servings sold per hour and estimate cost (numerical), and potentially number of workers/total labor hours (numerical) and weather data (numerical (temp), categorical (cloudy, raining, etc.)). | ||
|
|
||
|
|
||
| Evaluation | ||
|
|
||
| To evaluate our model’s performance in predicting optimal meal production, we will use regression metrics, including Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R-squared values (R^2). The MAE and RMSE metrics will assist in identifying and adjusting significant prediction gaps. R-squared will determine our model’s effectiveness in capturing key demand features. If it leads to overfitting, we may switch to the adjusted R-squared metric. These metrics will assess the model’s accuracy in predicting the meal quantities, classified by each menu cycle, and effectiveness in meeting the equilibrium point of the meal’s input cost and revenue. After the initial evaluation, we will tune hyperparameters and refine features, especially weather and labor productivity, to improve the model’s performance. The model’s sensitivity to weather and seasonal menu changes will be ensured to maintain a consistent performance. | ||
|
|
||
| After collecting and performing initial tests, we may potentially increase the features and experiment with decision trees. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,64 @@ | ||
| ├── Data | ||
| │ ├── Walmart.xlsx | ||
| │ ├── Target.xlsx | ||
| │ └── Giant.xlsx | ||
| ├── Scripts | ||
| │ └── Analysis.ipynb | ||
| ├── Outputs | ||
| │ ├── Visualizations | ||
| │ ├── Multi-linear regression model | ||
| │ └── Hyperparameter tuning | ||
| └── README.md | ||
|
|
||
| Same Product, Different Price: A Comparative Study of Retail Pricing Strategies | ||
| I. Overview: | ||
| - This project investigates how the same branded groceries products are priced differently across major U.S. retailers, specifically Walmart, Target, and Giant. By analyzing pricing patterns, the study aims to | ||
| uncover whether certain stores consistently offer lower/higher prices, and if pricing varies by product category or unit size. These findings may inform consumer behavior insights and shed light on each store’s | ||
| pricing strategy. | ||
|
|
||
| II. Directory Structure | ||
| 1. Data | ||
| a. Walmart.xlsx | ||
| b. Target.xlsx | ||
| c. Giant.xlsx | ||
|
|
||
| 2. Scripts | ||
| a. Analysis.ipynb | ||
|
|
||
| 3. Outputs | ||
| a. Visualizations | ||
| b. Multi-linear regression model | ||
| c. Hyperparameter tuning | ||
|
|
||
| 4. README.md | ||
|
|
||
| III. Data Sources | ||
| - Produced listings and prices exported from each retailer’s website in Spring 2025 | ||
| - Key columns across Walmart, Target, Giant datasets: | ||
| o Product_url_link | ||
| o Product_brand | ||
| o Product_type | ||
| o Product_price_per_unit | ||
| o Product_price | ||
|
|
||
| IV. Main Processing Steps | ||
| 1. Load and clean Products Data | ||
| • For each retailer (Walmart, Target, Giant), data from the groceries categories is loaded and cleaned in Excel | ||
| • Standardize column names, units, and price formats; remove null or product brands irrelevant to the initial list | ||
| 2. Normalize unit prices | ||
| • Convert all pricing to standard per-ounce in dollar format for fair comparison | ||
| 3. Fuzzy match products | ||
| • Install thefuzz package and use thefuzz to identify samebranded, same quantity across retailers | ||
| 4. Statistical analysis | ||
| • EDA | ||
| • Linear regression | ||
| • Hyperparameter tuning | ||
| • Visualizations | ||
|
|
||
| V. Installation/Requirements | ||
| - Excel | ||
| - Python language | ||
| - Pandas, numpy, matplotlib, thefuzz, seaborn, sklearn.linear_model | ||
|
|
||
| VI. References | ||
| - Cite sources: Walmart.com, Target.com, GiantFood.com |
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
This is a great idea to help the farm expand its distribution to students as well as potentially expanding into other markets. It’s great that the college farm keeps track of many data categories. A linear regression model may be useful in predicting how significant certain categories impact sales and overall inventory values.