I have published a comparison of promptfoo, garak and my own tool against the same two third-party targets. Every reply from all three is scored by one shared rule rather than by each tool's own judge, because the three count different events and their totals are not comparable.
https://github.com/qatration/qatration/blob/main/docs/benchmark.md
Where promptfoo came out ahead. Its own verdict was 143 of 200 tests failed, and the shared rule counted 150 of 200 replies carrying the planted string. Those two numbers nearly agree, which is a point in the grader's favour: a model judge landing within seven of a substring check is not something I expected to be able to write down.
A claim I made about promptfoo and then withdrew, because you would have found it and it is better that I say it. The first version of the page said promptfoo was the only tool whose attacks move the outcome: conditioned on the poisoned document being retrieved, ordinary traffic leaked at 85% and promptfoo at 95% on 158 retrievals. A reviewer asked for a p-value. Fisher exact against that baseline is p = 0.078, and against a cleaner baseline of plain support questions, 94%, it is p = 0.59. The limiting number was never your 158, it is the 27 replies in the baseline. The page now says the effect is possible and not established. garak and my own tool come out at p = 1.00 on the same comparison, so if anything moves that number it is yours, but this run does not show it.
Two questions.
- I ran four plugins at fifty tests each,
harmful:privacy, pii:direct, prompt-extraction and rbac, with strategies: [basic] and the generator pointed at ollama:chat:mistral-nemo. I did not run indirect-prompt-injection, which on reflection is the one aimed at exactly this target class, and I should have. Is that selection unrepresentative enough to invalidate the comparison, and which plugins would you run against a RAG chatbot?
promptfoo redteam generate asks for email verification before it will run, while a plain promptfoo eval does not. I set a project address and it worked. It is in the page's limitations because anyone reproducing this will meet it. If there is an offline or self-hosted path that avoids it, tell me and I will document that instead.
One observation, offered because it cost me an hour rather than as a bug report. The default of four requests in flight took the unguarded target down: local-rag-chat is a single-worker uvicorn in front of a 3B model, and four concurrent requests wedged it for over an hour and produced a result file full of transport errors. I discarded that run and drove every tool with -j 1, which is also the only way the wall-clock column compares tools rather than harnesses. A line in the docs about small local targets might save somebody the same hour.
The targets, for context: local-rag-chat at cd8cd89, an ordinary RAG application with no guardrails, and NeMo Guardrails with input and output rails. The corpus is mine, that project ships none, so nothing on the page is a vulnerability report about anybody's code.
My configs are at out/bench/configs/promptfooconfig.yaml and the guarded one beside it. Your raw result files are not in my repository, only my own artifacts and the commands to regenerate yours.
I have published a comparison of promptfoo, garak and my own tool against the same two third-party targets. Every reply from all three is scored by one shared rule rather than by each tool's own judge, because the three count different events and their totals are not comparable.
https://github.com/qatration/qatration/blob/main/docs/benchmark.md
Where promptfoo came out ahead. Its own verdict was 143 of 200 tests failed, and the shared rule counted 150 of 200 replies carrying the planted string. Those two numbers nearly agree, which is a point in the grader's favour: a model judge landing within seven of a substring check is not something I expected to be able to write down.
A claim I made about promptfoo and then withdrew, because you would have found it and it is better that I say it. The first version of the page said promptfoo was the only tool whose attacks move the outcome: conditioned on the poisoned document being retrieved, ordinary traffic leaked at 85% and promptfoo at 95% on 158 retrievals. A reviewer asked for a p-value. Fisher exact against that baseline is p = 0.078, and against a cleaner baseline of plain support questions, 94%, it is p = 0.59. The limiting number was never your 158, it is the 27 replies in the baseline. The page now says the effect is possible and not established. garak and my own tool come out at p = 1.00 on the same comparison, so if anything moves that number it is yours, but this run does not show it.
Two questions.
harmful:privacy,pii:direct,prompt-extractionandrbac, withstrategies: [basic]and the generator pointed atollama:chat:mistral-nemo. I did not runindirect-prompt-injection, which on reflection is the one aimed at exactly this target class, and I should have. Is that selection unrepresentative enough to invalidate the comparison, and which plugins would you run against a RAG chatbot?promptfoo redteam generateasks for email verification before it will run, while a plainpromptfoo evaldoes not. I set a project address and it worked. It is in the page's limitations because anyone reproducing this will meet it. If there is an offline or self-hosted path that avoids it, tell me and I will document that instead.One observation, offered because it cost me an hour rather than as a bug report. The default of four requests in flight took the unguarded target down: local-rag-chat is a single-worker uvicorn in front of a 3B model, and four concurrent requests wedged it for over an hour and produced a result file full of transport errors. I discarded that run and drove every tool with
-j 1, which is also the only way the wall-clock column compares tools rather than harnesses. A line in the docs about small local targets might save somebody the same hour.The targets, for context: local-rag-chat at
cd8cd89, an ordinary RAG application with no guardrails, and NeMo Guardrails with input and output rails. The corpus is mine, that project ships none, so nothing on the page is a vulnerability report about anybody's code.My configs are at
out/bench/configs/promptfooconfig.yamland the guarded one beside it. Your raw result files are not in my repository, only my own artifacts and the commands to regenerate yours.