Pricing: A free starter plan offers limited storage and one database. Pro plan starting at $49 per month offers many more traces, larger storage, and access to features such as saved workflows.

Standout feature: Real-time search for monitoring production environments at scale

Best suited for: Enterprise with larger challenges with substantial RAG integration

Promptfoo

LLMs can fail in a number of ways. Promptfoo iterates through various tests that simulate real user interactions to simulate the types of issues an LLM might face each day. Promptfoo also focuses on some of the biggest security problems and specializes in red teaming to detect any failure points that might be exposed by a malicious user. From toxic edge states to personally identifiable information (PII) leaks, the goal is to deliver tests that will expose potential jailbreaks and failures in the guardrails.

Pricing: “Free forever” means an open-source tool with community-based support. An enterprise version offers custom deployment options and better support.

Standout feature: Automated red-teaming and prompt scrutiny helps lock down implementations.

Best suited for: Security-focused teams that are constantly evaluating and re-evaluating their product’s security.

RAGAS

When RAG databases are a key part of the agentic stack, developers turn to RAGAS to stress test the deeper mathematical corners of the retrieval mechanism. The Python library offers standard and custom metrics for evaluating the performance of the RAG storage-and-retrieval mechanism at the level of vector mathematics. These measure behaviors such as faithfulness, relevance, and totality of recall. The philosophy begins with experiments to speed development but ends with fast integration with the deployment pipeline. Instead of just doing a “vibe check” on the RAG database, developers are using a more scientific approach to test and converge on better total performance.

Pricing: Fully open source under Apache 2.0 license

Standout feature: RAG focus helps teams relying on vector databases for knowledge curation.

Best suited for: Teams with a substantial reliance on RAG databases

Rhesis AI

Many of tools in this evolving market niche are designed for hard-core developers. Rhesis AI wants to bring other stakeholders into the development cycle so they can create tests and evaluate performance, too. That means domain experts, product managers, and even C-suite suits can track how the LLMs behave in conversations. Adversarial or confrontational engagements that devolve into the edge cases that bring headaches are easy to simulate repeatedly to optimize responses. The platform is designed to test all stages of development in a way that’s accessible to all stakeholders.

Pricing: Said to be “open source first” but with enterprise plans for those that need it.

Standout feature: The focus on putting humans in the loop is ideal for applications that require input from meat-based intelligence.

Best suited for: Applications requiring more collaboration with domain experts

Vellum

Anyone who needs a personal assistant can turn to Vellum to help build one that is trained on your data. Along the way, you will evaluate performance using its elaborate testing framework that tracks performance against any of the metrics and use cases you supply. Real-time dashboards track performance using metrics such as token usage costs, latency, or response quality. Multiple teams can work in parallel with version controls that allow iteration and competition. The end result is an agent that’s tuned to your needs.

Pricing: A basic free tier for experimentation. The Mighty starts at $30 per month and comes with more storage and compute credits.

Standout feature: End-to-end integration simplifies managing new development.

Best suited for: Cross-functional teams looking for a centralized solution with wide integration