## Problems most (all?) existing comparisons are purely quantitative (perplexity scores) ## Objectives - qualitative comparison (prompt input & outputs) ## See Also - https://arxiv.org/abs/2212.09720 - https://github.com/Hannibal046/Awesome-LLM - [https://tatsu-lab.github.io/alpaca_eval](https://tatsu-lab.github.io/alpaca_eval/) - https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard - https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard - https://paperswithcode.com/task/language-modelling - https://paperswithcode.com/sota/code-generation-on-humaneval
Problems
most (all?) existing comparisons are purely quantitative (perplexity scores)
Objectives
See Also