This project evaluates the performance and correctness of LLM-generated code compared to human-written implementations.
- Measure runtime performance
- Analyze memory usage
- Validate correctness on algorithmic problems
Implemented automated benchmarking and testing pipelines using Python.
- Python
- Large Language Models
- Algorithms
- Performance Analysis
Academic / Research Project