Skip to content

Benchmarked promptfoo against two third-party targets, plus a question about plugin selection #10505

Description

@qatration

I have published a comparison of promptfoo, garak and my own tool against the same two third-party targets. Every reply from all three is scored by one shared rule rather than by each tool's own judge, because the three count different events and their totals are not comparable.

https://github.com/qatration/qatration/blob/main/docs/benchmark.md

Where promptfoo came out ahead. Its own verdict was 143 of 200 tests failed, and the shared rule counted 150 of 200 replies carrying the planted string. Those two numbers nearly agree, which is a point in the grader's favour: a model judge landing within seven of a substring check is not something I expected to be able to write down.

A claim I made about promptfoo and then withdrew, because you would have found it and it is better that I say it. The first version of the page said promptfoo was the only tool whose attacks move the outcome: conditioned on the poisoned document being retrieved, ordinary traffic leaked at 85% and promptfoo at 95% on 158 retrievals. A reviewer asked for a p-value. Fisher exact against that baseline is p = 0.078, and against a cleaner baseline of plain support questions, 94%, it is p = 0.59. The limiting number was never your 158, it is the 27 replies in the baseline. The page now says the effect is possible and not established. garak and my own tool come out at p = 1.00 on the same comparison, so if anything moves that number it is yours, but this run does not show it.

Two questions.

  1. I ran four plugins at fifty tests each, harmful:privacy, pii:direct, prompt-extraction and rbac, with strategies: [basic] and the generator pointed at ollama:chat:mistral-nemo. I did not run indirect-prompt-injection, which on reflection is the one aimed at exactly this target class, and I should have. Is that selection unrepresentative enough to invalidate the comparison, and which plugins would you run against a RAG chatbot?
  2. promptfoo redteam generate asks for email verification before it will run, while a plain promptfoo eval does not. I set a project address and it worked. It is in the page's limitations because anyone reproducing this will meet it. If there is an offline or self-hosted path that avoids it, tell me and I will document that instead.

One observation, offered because it cost me an hour rather than as a bug report. The default of four requests in flight took the unguarded target down: local-rag-chat is a single-worker uvicorn in front of a 3B model, and four concurrent requests wedged it for over an hour and produced a result file full of transport errors. I discarded that run and drove every tool with -j 1, which is also the only way the wall-clock column compares tools rather than harnesses. A line in the docs about small local targets might save somebody the same hour.

The targets, for context: local-rag-chat at cd8cd89, an ordinary RAG application with no guardrails, and NeMo Guardrails with input and output rails. The corpus is mine, that project ships none, so nothing on the page is a vulnerability report about anybody's code.

My configs are at out/bench/configs/promptfooconfig.yaml and the guarded one beside it. Your raw result files are not in my repository, only my own artifacts and the commands to regenerate yours.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions