Skip to content

Add LLM-as-Judge task success evaluation to optimizer - #4

Open
MahtabSarvmaili wants to merge 22 commits into
mainfrom
mahtab_prompt_optimize
Open

Add LLM-as-Judge task success evaluation to optimizer#4
MahtabSarvmaili wants to merge 22 commits into
mainfrom
mahtab_prompt_optimize

Conversation

@MahtabSarvmaili

Copy link
Copy Markdown
Contributor

Implement LLMJudgeSuccess evaluator for automatic task completion assessment Add server-specific optimization with automatic server detection Enhanced documentation with evaluation criteria and output format

@MahtabSarvmaili
MahtabSarvmaili requested a review from saqadri July 22, 2025 03:41
Bill Peterson added 17 commits July 28, 2025 11:08
- evaluting the "is_sucessful" based on the LLM as Judge output

- removing the corret tool

-  removing the limitation on the number of successful and failed samples for optimization
- extracting the user's queries from the trace file
- Updated response extraction

- Updated core_trace_process.py to use new extraction functions

- Added helper functions for detecting different span types
- sorting the successful and failed examples based on the scores
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant