diff --git a/.DS_Store b/.DS_Store new file mode 100644 index 0000000..f706a12 Binary files /dev/null and b/.DS_Store differ diff --git a/.Rhistory b/.Rhistory new file mode 100644 index 0000000..e69de29 diff --git a/Idea 2.md b/Idea 2.md new file mode 100644 index 0000000..7f86f82 --- /dev/null +++ b/Idea 2.md @@ -0,0 +1,23 @@ +Christine Gurek +Data 400 +Professor Bilen +September 18, 2024 + + Idea 2 +1) Research question- try to be as focused and as specific as possible. The Effect book is an invaluable resource here. Take a look at the chapter readings from it on the syllabus. + +Does the increase of the minimum wage impact the amount of hours people are working in the retail sector? As more employees fight to raise the minimum wage, what may be the impact on the hours they are given? Many employees today are turning away from the average 9-5 and looking for more flexible hours, but can we have both higher wages and these types of flexibility, while also paying the bills? + +2) Data source, with details about how you will get the data, what variables the dataset has or you will create etc. This can get pretty long, which is fine! + +The Bureau of Labor Statistics has many data sets and visualizations available for people to use, I would get my data from here. It will be a fairly simple data set, exploring the average hours worked per week by those employees in the retail sector. The national minimum wage is easily trackable, but many companies are opting to give their employees higher minimum wages, so we can also use the Bureau of Labor Statistics information on that. The data can be broken down by year, month, and week. + +3) What you will do with the data, EDA, model specification. + +In terms of EDA I will look for any missing values or outliers that may be impacting the outcome, as well as finding a minimum and maximum value, as this may give some insight into seasonality, or help to visualize the trends that have arisen in the past years. +I will use a time series model in order to predict what the number of hours worked per week, on average will be as the minimum wages paid increases, and therefore see if there will be any further changes. + +4) What are the ethical implications of your project? + +The ethical implications of my project are more on the positive side, it can help employees look at the amount that they are being paid currently and their own amount of weekly hours on average, and use it to negotiate a better wage, or a better work schedule. Another implication of my project is that it could show a decrease in hours as the wage increases, this could force employees to find second or third jobs, and eventually turn towards burnout as so many are. But is it really fair to keep some employees at only $7.25 per hour? It is a very tough question to answer, but I hope this research could bring us towards a better solution for everyone. + diff --git a/Idea 3.md b/Idea 3.md new file mode 100644 index 0000000..d140ed9 --- /dev/null +++ b/Idea 3.md @@ -0,0 +1,22 @@ +Christine Gurek +Data 400 +Professor Bilen +September 18, 2024 + + Idea 3 +1) Research question- try to be as focused and as specific as possible. The Effect book is an invaluable resource here. Take a look at the chapter readings from it on the syllabus. + +Does a Video Game’s sales value reflect the likeability of the game? In other words, do games that generate more sales, also have better reviews than those that do not sell as well. This could also be further investigated by which company published the game, and the different resources provided by these publishers. If the publisher has more revenue, there is larger teams working on the games and more marketing being done, which usually results in higher sales. + +2) Data source, with details about how you will get the data, what variables the dataset has or you will create etc. This can get pretty long, which is fine! + +There is a dataset on Kaggle that has a list of 16600 video games with their names, Platform, Year it was released, Genre, Publisher, and sales data. I was able to make an account and download the csv file from there. There are also websites where you are able to data-scrape the reviews, this however may take longer to complete than is allotted for this project. I would also then have to run a text analysis on the reviews, unless they are already given a numerical rating. + +3) What you will do with the data, EDA, model specification. + +I will check for duplicates or prequels, for example Fifa 14-16, which may affect the overall sales of the games. I will also check for any outliers or missing data that may impact the outcome of my research. I will use a regression model in order to explore the relationship between the sales and the reviews surrounding the video game. + +4) What are the ethical implications of your project? + +This project would be fairly ethical, it would not affect many people directly, but it may open the door to more research on how to make better games, focusing on the types of games we know people enjoy. It does bring up the question though of how do we create better games without also creating a monopoly within the gaming industry? for example if you need a bigger programming team and more marketing, only people with a certain amount of money will be able to achieve it, so how do we make it more accessible to create video games while also creating better games. + diff --git a/Mini Project Idea.md b/Mini Project Idea.md new file mode 100644 index 0000000..458f867 --- /dev/null +++ b/Mini Project Idea.md @@ -0,0 +1,25 @@ + +Christine Gurek +Data 400 +Professor Bilen +September 10, 2024 +Mini-Project Idea +1) Research question- try to be as focused and as specific as possible. The Effect book is an invaluable resource here. Take a look at the chapter readings from it on the syllabus. + +At what age are dancers most vulnerable to injury? When should injury prevention education begin for dancers? +Many dancers will begin classes at a very early age, oftentimes starting at 2-3 years old, and will continue dancing throughout their lives in some way. Research shows that dancers will average one injury per year, but when are the most intense and long-lasting injuries occurring? Since smaller injuries are more normalized for dancers, many times they will not seek medical attention and will attempt to heal themselves at home. I think it’s very important to educate young dancers on injury prevention and recovery, but when is the time to begin this education? I had no form of this education until college and by then I had already experienced multiple injuries, many I’m sure will be long lasting as they were not treated properly, so this is a very important topic to me personally as well as many dance companies, because if many dancers are injured in their young lives they may not be able to perform to the best of their abilities as they continue on in their career, which is not great for companies trying to hire long-term dancers. This education will not only help companies find long term dancers but will allow dancers who want to pursue this type of career a better chance at achieving that goal, many dancers will experience career ending injuries at a very young age, so having this type of education will allow them to continue exploring their passion as a dancer in a safer way. This also connects directly to Dickinson as we offer a class on anatomy and injury prevention, as well as have a seminar on injury prevention within the Dance Theater Group, but we do not have access to the same resources as the Athletics Department. This could also help bring change to the local dance studios in the Carlisle area if Dickinson can find a way to teach them about injury prevention as well, especially if my findings do find a younger age to be the most vulnerable to more intense dance-related injuries. + +2) Data source, with details about how you will get the data, what variables the dataset has or you will create etc. This can get pretty long, which is fine! + +I will be using Data from the NEISS database, it is a free database from participating Emergency Rooms around the United States. It compiles data about each patient, what they were seen for, what their diagnosis was, etc. I was able to submit a query on the website to get the information I wanted from the database: dance related injuries on people ages 3-40 years old, all genders, all races, what their diagnosis was, and what they originally came in for from the past 10 years, so 2014-2024. After submitting this query the NEISS gave me an excel sheet that I was able to download and will further clean up from there. The data has been broken down by years, so once I have cleaned up each individual year I will put them together for my analysis. + +3) What you will do with the data, EDA, model specification. + +There is a large amount of data available to me so I may not include all 10 years, I also plan to narrow down the data to those who are active dancers, whether it is just classes, or a professional career, so that the people who were just dancing with friends and tripped are not included. Once I have cleaned up the data I plan to do some EDA to find any outliers that may affect my results, find if there is anything missing from the columns I am hoping to use, and look at the shape of the age data as well to better decide on a model, or if I need to do any further cleaning/data manipulation. I will then sort the data by age and figure out the most frequent age that comes up. Once I have found that I also plan to find the most intense injuries/diagnoses and find which age they are mostly occurring at, the injury type data has already been converted into a numerical format so this will be helpful in this process. For the ages I can use a bar plot to better visualize the ages that are present, and I will use a regression model for the intensity of the injury vs. the age of the patient to better determine when the best time for dancers’ injury prevention education is to begin. + + + + + + + diff --git a/Mini-Project Presentation.pptx b/Mini-Project Presentation.pptx new file mode 100644 index 0000000..acc0ed0 Binary files /dev/null and b/Mini-Project Presentation.pptx differ diff --git a/README copy.md b/README copy.md new file mode 100644 index 0000000..ed45ff5 --- /dev/null +++ b/README copy.md @@ -0,0 +1,16 @@ +Ballerina Diagram + + +## Dance Injury Analysis +### By Christine Gurek and Heidi Beardsley +Our project aims to look into the types of injuries that dancers experience, where these injuries are most likely to happen, how they are happening, and if age can predict these injuries. Our main aim is to investigate when injury prevention education should begin for dancers, so that the number of long-term injuries can be prevented. + +We did this by using data from the National Electronic Injury Surveillance System. On their website you are able to submit a query with the defined parameters of data you need, which for this project was dance injuries from the past 10 years. After submitting this query, we were able to download the excel files to our computers and begin the data cleaning and analysis process. + +For data cleaning we first took out any unnecessary columns. Then we filtered the narrative column to rule out any rows that did not include the words “dance” and “class” in them, so we could be sure we were only looking into the injuries of long-term dancers. We then imported the data into Jupyter Notebook and concatenated all of them together into one dataframe. The data now included 6 columns and 4911 rows. + +Then we completed our Exploratory Data Analysis. From this we were able to find that the average age of the dancers in the sample was 14, and the most common injuries were a strain/sprain. We also looked into the relationships between the months and the frequencies of injuries, where we found that February and October were the months with the most frequent injuries happening. We also visualized the proportion of injuries onto a photo of a dancer in order to show where these injuries were most often occurring, which ended up being the knee and ankle. We also ran a Latent Dirichlet Analysis on the narrative column to see if there was a pattern in the way dancers were getting injured. + +We then ran Random Forrest Classifier and SVM Classifier models on our data to first predict whether a dancer experienced a strain or not, and then if they experienced a fracture or not. Both models did fairly well, but the fracture models did better in terms of accuracy. We also found the feature importances for each column in contributing to the models and in both the sprain and fracture models, the body part was the most important feature for prediction. + +Through our results, we uncovered several important findings related to injury patterns in dancers. We found that dancers are most vulnerable to injury between the ages of 10 and 15, with the highest incidence occurring at age 14, making this age group particularly prone to serious injuries such as sprains, strains, and fractures. Injury prevention education should begin early, ideally before age 10, as injury rates start to rise at this point. Early education on which body parts are most prone to injury, the types of injuries that can occur, and how seasonal patterns affect injury risk can help dancers make more informed decisions about their training and warm-ups. Additionally, our models confirmed that age is a strong predictor of sprains, strains, and fractures, but body part was the most critical factor in predicting these injuries, highlighting the need for targeted education on specific injury-prone areas like the knee, ankle, and foot. The most common injuries, sprains, strains, and fractures, typically occur in the knee, ankle, and foot, and our analysis revealed that terms like “fell,” “pop,” and “twisted” were commonly associated with these injuries, offering insights into their causes. In conclusion, our findings stress the importance of early injury prevention education focused on vulnerable age groups and injury-prone body parts, helping dancers reduce their injury risk and make safer, smarter choices in their training. diff --git a/Untitled Document.md b/Untitled Document.md new file mode 100644 index 0000000..29ac0fe --- /dev/null +++ b/Untitled Document.md @@ -0,0 +1,192 @@ +# Dillinger +## _The Last Markdown Editor, Ever_ + +[![N|Solid](https://cldup.com/dTxpPi9lDf.thumb.png)](https://nodesource.com/products/nsolid) + +[![Build Status](https://travis-ci.org/joemccann/dillinger.svg?branch=master)](https://travis-ci.org/joemccann/dillinger) + +Dillinger is a cloud-enabled, mobile-ready, offline-storage compatible, +AngularJS-powered HTML5 Markdown editor. + +- Type some Markdown on the left +- See HTML in the right +- ✨Magic ✨ + +## Features + +- Import a HTML file and watch it magically convert to Markdown +- Drag and drop images (requires your Dropbox account be linked) +- Import and save files from GitHub, Dropbox, Google Drive and One Drive +- Drag and drop markdown and HTML files into Dillinger +- Export documents as Markdown, HTML and PDF + +Markdown is a lightweight markup language based on the formatting conventions +that people naturally use in email. +As [John Gruber] writes on the [Markdown site][df1] + +> The overriding design goal for Markdown's +> formatting syntax is to make it as readable +> as possible. The idea is that a +> Markdown-formatted document should be +> publishable as-is, as plain text, without +> looking like it's been marked up with tags +> or formatting instructions. + +This text you see here is *actually- written in Markdown! To get a feel +for Markdown's syntax, type some text into the left window and +watch the results in the right. + +## Tech + +Dillinger uses a number of open source projects to work properly: + +- [AngularJS] - HTML enhanced for web apps! +- [Ace Editor] - awesome web-based text editor +- [markdown-it] - Markdown parser done right. Fast and easy to extend. +- [Twitter Bootstrap] - great UI boilerplate for modern web apps +- [node.js] - evented I/O for the backend +- [Express] - fast node.js network app framework [@tjholowaychuk] +- [Gulp] - the streaming build system +- [Breakdance](https://breakdance.github.io/breakdance/) - HTML +to Markdown converter +- [jQuery] - duh + +And of course Dillinger itself is open source with a [public repository][dill] + on GitHub. + +## Installation + +Dillinger requires [Node.js](https://nodejs.org/) v10+ to run. + +Install the dependencies and devDependencies and start the server. + +```sh +cd dillinger +npm i +node app +``` + +For production environments... + +```sh +npm install --production +NODE_ENV=production node app +``` + +## Plugins + +Dillinger is currently extended with the following plugins. +Instructions on how to use them in your own application are linked below. + +| Plugin | README | +| ------ | ------ | +| Dropbox | [plugins/dropbox/README.md][PlDb] | +| GitHub | [plugins/github/README.md][PlGh] | +| Google Drive | [plugins/googledrive/README.md][PlGd] | +| OneDrive | [plugins/onedrive/README.md][PlOd] | +| Medium | [plugins/medium/README.md][PlMe] | +| Google Analytics | [plugins/googleanalytics/README.md][PlGa] | + +## Development + +Want to contribute? Great! + +Dillinger uses Gulp + Webpack for fast developing. +Make a change in your file and instantaneously see your updates! + +Open your favorite Terminal and run these commands. + +First Tab: + +```sh +node app +``` + +Second Tab: + +```sh +gulp watch +``` + +(optional) Third: + +```sh +karma test +``` + +#### Building for source + +For production release: + +```sh +gulp build --prod +``` + +Generating pre-built zip archives for distribution: + +```sh +gulp build dist --prod +``` + +## Docker + +Dillinger is very easy to install and deploy in a Docker container. + +By default, the Docker will expose port 8080, so change this within the +Dockerfile if necessary. When ready, simply use the Dockerfile to +build the image. + +```sh +cd dillinger +docker build -t /dillinger:${package.json.version} . +``` + +This will create the dillinger image and pull in the necessary dependencies. +Be sure to swap out `${package.json.version}` with the actual +version of Dillinger. + +Once done, run the Docker image and map the port to whatever you wish on +your host. In this example, we simply map port 8000 of the host to +port 8080 of the Docker (or whatever port was exposed in the Dockerfile): + +```sh +docker run -d -p 8000:8080 --restart=always --cap-add=SYS_ADMIN --name=dillinger /dillinger:${package.json.version} +``` + +> Note: `--capt-add=SYS-ADMIN` is required for PDF rendering. + +Verify the deployment by navigating to your server address in +your preferred browser. + +```sh +127.0.0.1:8000 +``` + +## License + +MIT + +**Free Software, Hell Yeah!** + +[//]: # (These are reference links used in the body of this note and get stripped out when the markdown processor does its job. There is no need to format nicely because it shouldn't be seen. Thanks SO - http://stackoverflow.com/questions/4823468/store-comments-in-markdown-syntax) + + [dill]: + [git-repo-url]: + [john gruber]: + [df1]: + [markdown-it]: + [Ace Editor]: + [node.js]: + [Twitter Bootstrap]: + [jQuery]: + [@tjholowaychuk]: + [express]: + [AngularJS]: + [Gulp]: + + [PlDb]: + [PlGh]: + [PlGd]: + [PlOd]: + [PlMe]: + [PlGa]: diff --git a/original_ data/NEISS_2014.XLSX b/original_ data/NEISS_2014.XLSX new file mode 100644 index 0000000..5d598ac Binary files /dev/null and b/original_ data/NEISS_2014.XLSX differ diff --git a/original_ data/NEISS_2015.XLSX b/original_ data/NEISS_2015.XLSX new file mode 100644 index 0000000..0d19dea Binary files /dev/null and b/original_ data/NEISS_2015.XLSX differ diff --git a/original_ data/NEISS_2016.XLSX b/original_ data/NEISS_2016.XLSX new file mode 100644 index 0000000..0ff3b3d Binary files /dev/null and b/original_ data/NEISS_2016.XLSX differ diff --git a/original_ data/NEISS_2017.XLSX b/original_ data/NEISS_2017.XLSX new file mode 100644 index 0000000..b4b2d5f Binary files /dev/null and b/original_ data/NEISS_2017.XLSX differ diff --git a/original_ data/NEISS_2018.XLSX b/original_ data/NEISS_2018.XLSX new file mode 100644 index 0000000..546e53c Binary files /dev/null and b/original_ data/NEISS_2018.XLSX differ diff --git a/original_ data/NEISS_2019.XLSX b/original_ data/NEISS_2019.XLSX new file mode 100644 index 0000000..89c254e Binary files /dev/null and b/original_ data/NEISS_2019.XLSX differ diff --git a/original_ data/NEISS_2020.XLSX b/original_ data/NEISS_2020.XLSX new file mode 100644 index 0000000..eb74921 Binary files /dev/null and b/original_ data/NEISS_2020.XLSX differ diff --git a/original_ data/NEISS_2021.XLSX b/original_ data/NEISS_2021.XLSX new file mode 100644 index 0000000..f0b0fcf Binary files /dev/null and b/original_ data/NEISS_2021.XLSX differ diff --git a/original_ data/NEISS_2022.XLSX b/original_ data/NEISS_2022.XLSX new file mode 100644 index 0000000..a1b29e8 Binary files /dev/null and b/original_ data/NEISS_2022.XLSX differ diff --git a/original_ data/NEISS_2023.XLSX b/original_ data/NEISS_2023.XLSX new file mode 100644 index 0000000..c5f2edb Binary files /dev/null and b/original_ data/NEISS_2023.XLSX differ diff --git a/original_ data/NEISS_FMT.XLSX b/original_ data/NEISS_FMT.XLSX new file mode 100644 index 0000000..0973068 Binary files /dev/null and b/original_ data/NEISS_FMT.XLSX differ