Holistic Evaluation of Generative AI Systems // Jineet Doshi // #280

27 snips

Dec 23, 2024

In this insightful discussion, Jineet Doshi, an award-winning AI lead with over seven years at Intuit, dives deep into the complexities of evaluating generative AI systems. He emphasizes the importance of holistic evaluation to foster trust and the unique challenges posed by large language models. Jineet explores diverse evaluation methods, from classic NLP techniques to innovative strategies like red teaming. He also tackles the financial nuances of generative AI and the balance between human insight and automated feedback for robust assessments.

Ask episode

AI Snips

Chapters

Transcript

Episode notes

INSIGHT

LLM Evaluation Challenges

Evaluating LLMs is challenging due to their open-ended outputs and broad capabilities.
Traditional ML metrics are inadequate for assessing nuanced tasks like poem generation.

INSIGHT

Traditional NLP Techniques for LLM Evaluation

Traditional NLP techniques can be applied to LLM evaluation by using multiple-choice questions or text similarity.
However, these methods have limitations in evaluating open-ended tasks and can be sensitive to the choice of embedding models.

ADVICE

Using Benchmarks for LLM Evaluation

Use benchmarks to evaluate LLMs across various factors like knowledge, reasoning, and toxicity.
Be mindful of benchmark limitations, data leakage, and the need for custom benchmarks for specific use cases.

Get the Snipd Podcast app to discover more snips from this episode

Get the app

Jineet Doshi is an award-winning Scientist, Machine Learning Engineer, and Leader at Intuit with over 7 years of experience. He has a proven track record of leading successful AI projects and building machine-learning models from design to production across various domains, which have impacted 100 million customers and significantly improved business metrics, leading to millions of dollars of impact.

Holistic Evaluation of Generative AI Systems // MLOps Podcast #280 with Jineet Doshi, Staff AI Scientist or AI Lead at Intuit.

// Abstract

Evaluating LLMs is essential in establishing trust before deploying them to production. Even post-deployment, evaluation is essential to ensure LLM outputs meet expectations, making it a foundational part of LLMOps. However, evaluating LLMs remains an open problem. Unlike traditional machine learning models, LLMs can perform a wide variety of tasks, such as writing poems, Q&A, summarization, etc. This leads to the question of how do you evaluate a system with such broad intelligence capabilities? This talk covers the various approaches for evaluating LLMs, such as classic NLP techniques, red teaming, and newer ones like using LLMs as a judge, along with the pros and cons of each. The talk includes an evaluation of complex GenAI systems like RAG and Agents. It also covers evaluating LLMs for safety and security, and the need to have a holistic approach for evaluating these very capable models.

// Bio

Jineet Doshi is an award-winning AI Lead and Engineer with over 7 years of experience. He has a proven track record of leading successful AI projects and building machine learning models from design to production across various domains, which have impacted millions of customers and have significantly improved business metrics, leading to millions of dollars of impact. He is currently an AI Lead at Intuit, where he is one of the architects and developers of their Generative AI platform, which is serving Generative AI experiences for more than 100 million customers around the world. Jineet is also a guest lecturer at Stanford University as part of their building LLM Applications class. He is on the Advisory Board of the University of San Francisco’s AI Program. He holds multiple patents in the field, is on the steering committee of MLOps World Conference, and has also co-chaired workshops at top AI conferences like KDD. He holds a Master's degree from Carnegie Mellon University.

// MLOps Swag/Merch

https://shop.mlops.community/

// Related Links

Website: https://www.intuit.com/

--------------- ✌️Connect With Us ✌️ -------------

Join our Slack community: https://go.mlops.community/slack

Catch all episodes, blogs, newsletters, and more: https://mlops.community/

Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/

Connect with Jineet on LinkedIn: https://www.linkedin.com/in/jineetdoshi/

Timestamps:[00:00] Jineet's preferred coffee[00:20] Takeaways[01:24] Please like, share, leave a review, and subscribe to our MLOps channels![01:36] LLM evaluation at scale[03:13] Challenges in GenAI evaluation[08:09] Eval products vs platforms[09:28] Evaluation methods for models[14:03] NLP evaluation techniques[25:06] LLM as a judge/jury[31:56] LLMs and pizza brainstorming[34:07] Cost per answer breakdown[38:29] Evaluating RAG systems[44:00] Testing with LLMs and humans[49:23] Evaluating AI use cases[54:19] AI workflow stress testing[55:40] Wrap up