Solutions
Discover how you can run your deals on Cobl, by use-case, industry or role.
All departments
Inside Cobl

Claude vs ChatGPT vs Cobl for RFP responses: 2026 benchmark

Which AI model writes the best RFP response? We scored 8 models on accuracy, completeness, issue spotting and 6 other criteria, then compared cost and delivery time. We also ran Astra 6 and Fable 5.1 directly in ChatGPT and Claude.ai: both scored 14 to 22 points lower than through Cobl.

We started running evals 1 month ago. It has become 1 of our most valuable product investments. The setup takes time, costs money and comes with a steep learning curve, but it shows us exactly where our AI succeeds and where it fails. So far, we have run more than 60 tests across 15 models, generated 3,500 pages and processed 458M tokens.

At Cobl, our customers use AI to prepare RFP responses: go/no-go dashboards, technical proposals, security questionnaires and more. These opportunities often exceed €1M. Buyers read and score the documents, and a fabricated claim, an unfinished section or a missed mandatory clause can disqualify a bid. Every week, we ask the same question: is this response ready to send? Our evals help us understand what needs to improve before the answer can be yes.

Key takeaways

  • Cobl tested 8 AI models on the same real public-sector RFP, with 7 scripted messages, 10 planted contradictions and a 100-point quality grid.
  • Astra 6 scored highest at 87/100, but its run cost $68.09, about 8.5 times more than Sol 5.6 at $8.05.
  • Cobl chose Sol 5.6 as its default model: 73/100, zero fabricated facts across the month's tests, and mid-range cost and delivery time.
  • The same model scored higher through Cobl: Astra 6 got 87/100 versus 65/100 in ChatGPT, and Fable 5.1 got 74/100 versus 60/100 in Claude.ai.
  • Terra 5.6 was the fastest model, delivering a shorter but accurate RFP response in 40 minutes.

Evals in 1 paragraph

An eval is a repeatable test. We give several AI models the same RFP, knowledge base and scripted user messages. We collect their documents, conversations and traces: the record of which sources they read and which skills and tools they used. We then assess them against a consistent set of rules. The result is a quality ranking, a cost and speed report, and a list of problems assigned to the people responsible for fixing them.

The test

We use a real public-sector RFP. It includes 3 lots, a 40-page specification, an evaluation grid and an annex with requirements that can disqualify a bid. We also use a real knowledge base with hundreds of documents in it.

Each model receives the same 7 messages in the same order. The conversation moves from the brief and architecture diagrams through corrections, technical and pricing documents, advice requests, permission to proceed and customer references. It should produce 3 deliverables: a technical memo, a financial proposal, and a references and governance document. We assess both the conversation and the final files.

The source material is deliberately messy. We planted 10 contradictions in the test files, covering inconsistent prices, conflicting quantities and gaps between the proposed infrastructure and the RFP's requirements. The knowledge base adds other challenges: conflicting certification claims, missing product information and standard contract wording that clashes with the RFP.

We want to see how the agent handles uncertainty. Does it compare sources, explain conflicts and ask for a decision when needed? Or does it quietly choose whichever answer seems most plausible? Finding a problem is only part of the job. The agent also needs to respect the user's authority and carry unresolved issues into the final documents.

The criteria

We score quality across 9 criteria, for a total of 100 points.

Heading Points What decides the score
Accuracy 15 Are factual claims and figures supported by the knowledge base, the RFP or the user messages? We also check for invented prices. Even a true fact from the model's training counts as fabrication in this test if the approved sources do not support it.
Context understanding 15 Did the agent read the RFP, diagrams and uploads before writing? Did it understand the scope, partner responsibilities and conflicting instructions, and retain that context across 7 turns and 3 documents? We check the traces to find out.
Completeness 15 Were all requested documents delivered, with the expected length, depth and supporting figures?
RFP requirements 15 Did the response follow the required structure and answer every mandatory commitment? This includes an explicit penalty table meeting the RFP's minimum requirements.
Quality of reasoning 10 How well did the agent explain its recommendations and answer the user's questions? Did it stay within the decisions the user had authorised it to make?
Issue spotting 10 Did it identify the 10 planted contradictions and any additional problems? An issue raised only in chat earns 50% credit; carrying it into the document earns full credit.
Writing quality 8 Is the writing clear, professional and in French throughout, as required for this RFP? Does it avoid unfinished text, internal notes and traces of the generation process?
Persuasion 7 Does the response make a convincing case using supported claims? Persuasion based on exaggerated or unsupported promises scores below the midpoint.
Final document quality 5 Are the structure and typography clean? Does the delivered file contain the work and corrections the agent said it had completed?

Cost and time sit alongside this quality score. We show them in separate charts: total response cost and recorded delivery time. We choose a model by looking at all 3 dimensions.

Where the models stand

We tested 8 models through Cobl, using the same 7-turn scenario and the same application version. We also ran the scenario directly in ChatGPT with Astra 6 and in Claude.ai with Fable 5.1. The quality ranking shows all 10 runs: 8 through Cobl and 2 directly in ChatGPT and Claude.ai. White bars with grey outlines identify the direct runs. Separate charts compare each model across the 2 environments, and hollow markers identify the direct runs in the cost and time charts.

Each plotted result comes from 1 graded run. These results show how the models behaved in this scenario; they are not a definitive ranking across every RFP.

Quality

Which model produced the strongest response? The scores reveal different strengths, but also different ways of failing.

Bar chart ranking 10 RFP response runs by quality score, from Astra 6 at 87/100 through Cobl to Gemini 3.8 Flash at 48/100.

The 4 highest-scoring models illustrate why we look beyond the overall score.

  • Astra 6 combines thorough analysis with strong source discipline. It reads extensively before drafting, checks the provided documents against the RFP and catches all 10 planted contradictions. It also identifies less obvious conflicts in schedules, responsibilities and reused content. Crucially, it explains the financial discrepancies in the final document and quantifies the impact of a technical shortfall. It provides an explicit penalty table and avoids importing unsupported facts, prices or certifications.

With an accuracy score of 14/15, it still has weaknesses: it repeatedly assigns a partner's responsibility to the bidder, and its cautious language weakens some commitments. Even so, its graded run was the 1st complete response this month to pass all our mandatory checks.

  • Opus 5 analyses problems better than it resolves them in the final response. It produces the deepest financial analysis of the month, identifies all 7 planted financial contradictions and catches an additional sovereignty risk. But those findings do not reliably shape what it delivers. The final response declares compliance despite a shortfall it has already identified, repeats standard penalty wording that conflicts with the RFP, and leaves the financial proposal in chat instead of exporting it. It also adds a vendor fact absent from the approved sources. Its main weakness is the gap between sound analysis and a consistent, complete deliverable.
  • Fable 5.1 is comprehensive, but takes decisions beyond its authority. It earns full marks for completeness and RFP requirements, with detailed coverage, supporting figures and an explicit penalty table. Its accuracy score is only 2/15. It treats unapproved customer references as cleared, substitutes public prices for the figures in the provided documents and adds costs based on its own assumptions. It explains these choices in chat, then implements them anyway. The problem is that it treats identifying a possible improvement as permission to change the offer. In these runs, both Fable and Opus sometimes favour their own knowledge or interpretation of the RFP over the approved sources.
  • Sol 5.6 offers the most balanced profile for our default use. It fabricates no facts, identifies conflicting source information and proposes clear, traceable financial corrections. It scores 73/100. Its main failures concern follow-through: a correction announced in chat does not reach every document, the references document is not exported, and the penalty commitment lacks the required table. Its reasoning is useful and grounded, but the delivery process needs stronger checks.
  • Terra 5.6 and Luna 5.6 stay accurate, but provide less depth. They score 15/15 and 14/15 on accuracy, with overall scores of 65/100 and 55/100. Their proposals reach roughly 1/3 of the target length, include few supporting figures and identify fewer issues. Terra accepts the penalty requirements without supplying the table; Luna defers the answer. Luna also repeats a currency-formatting problem. Their restraint reduces unsupported claims, but leaves more analysis and completion work to the user.
  • Sonnet 5 does not challenge the source material enough. It carries financial errors into the response, treats a discrepancy as rounding and confuses certification wording. It raises few substantive reservations, misses the penalty clause and does not export to Word. The central weakness is insufficient verification before presenting the answer.
  • Gemini 3.8 Flash is fast and confident, including when the sources do not justify confidence. It catches 0 planted contradictions and describes conflicting financial files as fully aligned. It adds unsupported claims about qualifications, data-loss guarantees and legal protection, including a qualification error that automatically fails our checks. It does answer the governance requirements, explicitly accept the penalty minimums and deliver all 3 documents in 42 minutes. But its polished, assertive writing masks uncertainty the user needs to see.

What does Cobl add to ChatGPT and Claude.ai?

Choosing a model answers only part of the question. If that model is already available in ChatGPT or Claude.ai, what does using it through Cobl add? We tested Astra 6 in Cobl and ChatGPT, and Fable 5.1 in Cobl and Claude.ai, using the same RFP, knowledge base, 7 scripted messages and 100-point scoring grid.

These comparisons assess the whole working environment around the model: the instructions, skills and tools that guide how it reads sources, checks requirements and delivers documents. The practical question is whether Cobl produces a more complete, usable RFP response, with fewer omissions and less work left to the user. The charts below show the overall scores; the analysis explains where Cobl helps and where it still falls short.

RFP Response with Astra 6

Chart comparing Astra 6 quality scores for the same RFP: 87/100 through Cobl versus 65/100 directly in ChatGPT.

Astra 6 scores 87/100 in Cobl and 65/100 in ChatGPT, a 22-point difference in this test. Both runs stay close to the approved sources, with accuracy scores of 14/15 and 13/15. The larger difference is how that knowledge becomes a finished response. In Cobl, Astra checks the requirements more thoroughly, carries identified problems into the documents and supplies the mandatory penalty table. In ChatGPT, it produces a shorter response, leaves that requirement unresolved and includes internal review language in the client document.

For the user, this means less completion and editing work after generation. The clearest gains are in RFP requirements, issue spotting and writing quality. The corrected Cobl run costs $68.09 and spans a 65-minute session window, compared with $25.37 and 67 recorded minutes in ChatGPT. Cobl produces the stronger response at a higher cost, with a similar recorded delivery time.

RFP Response with Fable 5.1

Chart comparing Fable 5.1 quality scores for the same RFP: 74/100 through Cobl versus 60/100 directly in Claude.ai.

Fable 5.1 scores 74/100 in Cobl and 60/100 in Claude.ai, a 14-point difference. In Cobl, it delivers a more complete response, supplies the penalty table and identifies financial contradictions that the direct run misses. Its issue-spotting score rises from 3/10 to 8/10. The direct run also leaves more cleanup for the user, including spreadsheet references in the prose and a table announced but never written.

The improvement is uneven. Fable's accuracy falls from 8/15 in Claude.ai to 2/15 in Cobl because it introduces unsupported prices, assumes consent and changes the offer beyond its authority. Cobl improves coverage and follow-through in this run, but does not make those decisions reliable. The more extensive response also takes 126 active minutes versus 34 in Claude.ai.

Together, these comparisons suggest that the workflow around a model affects the response it delivers. They also show why we inspect every criterion: a higher total can hide a serious weakness. Each comparison contains only 1 graded run per setup, so the gaps need repeated testing before we can treat them as stable.

Quality against total response cost

Chart of quality score against total RFP response cost per model, from $0.38 for Luna 5.6 to $68.09 for Astra 6.

The chart shows the total recorded cost of each response, without dividing it by the number of pages produced. Among the Cobl runs, costs range from $0.38 for Luna to $68.09 for Astra. Sol scores 73/100 at $8.05, compared with 79/100 at $17.10 for Opus and 74/100 at $21.82 for Fable. The question is how much we spend to obtain a response of the required quality, including the work still needed before it can be sent.

Luna's and Gemini's prices are provisional: they use our gateway's entry-tier rates and still need confirmation. A pricing correction would change their position on the cost axis, without changing their quality score.

Why Astra costs more. The corrected graded run costs $68.09, about 8.5 times Sol's $8.05. This is the cost of the plotted run, not a sum of multiple attempts.

  1. A substantial workload at a higher total cost. Astra's corrected run used 21.9M prompt tokens and 118,000 completion tokens across 177 LLM calls. Sonnet used 18.2M input tokens and 144,000 output tokens for $3.99. Astra's total bill is about 17 times higher. Token volume alone does not explain that gap; pricing and caching also affect the bill.
  2. Reading and checking are part of the workload. Astra repeatedly revisits the sources before and during drafting. That supports its strong context understanding, but also means processing context across many LLM calls. The 21.9M prompt tokens are cumulative input across the run, not the size of the knowledge base.
  3. Writing more is not the main explanation. Opus and Fable each generated nearly 3 times Astra's output tokens, while their total recorded costs were $17.10 and $21.82, versus $68.09 for Astra. A longer answer does not necessarily cost more.

The practical priorities are to preserve useful source checking, reduce repeated processing, keep the cache effective across turns and prevent avoidable restarts. Prompts alone do not change the model's price tier.

Quality against time to deliver

Chart of quality score against delivery time per model, from 40 minutes for Terra 5.6 to 126 minutes for Fable 5.1.

Among the Cobl runs, recorded delivery time ranges from 40 minutes for Terra to 126 minutes for Fable. Astra's corrected session window is 65 minutes. Its highest quality score therefore comes with a higher cost, but not the longest delivery time. The direct runs record 67 minutes in ChatGPT and 34 minutes in Claude.ai. Astra's 65 minutes cover the session window; the other figures retain their recorded active durations.

The workflow also matters. Waiting for unnecessary approval or failing to export a completed document can prevent a usable delivery, regardless of how well the model writes. We therefore look at speed alongside depth, completeness and successful export.

Conclusion

Astra 6 delivers the strongest quality result in this test, followed by Opus 5 and Fable 5.1. But their strengths come with different trade-offs. Astra combines thorough analysis with disciplined use of sources, at a much higher cost. Opus reasons deeply but does not consistently apply its findings to the final files. Fable produces comprehensive responses but makes unsupported assumptions and unauthorised changes.

We have selected Sol 5.6 as our default. It combines strong overall quality, 0 fabricated facts across the month's tests, and mid-range cost and delivery time. Its observed weaknesses, such as incomplete exports and corrections not reaching every document, give us clear targets for improving Cobl's skills, prompts and tools. That does not make every failure easy to fix, but it gives us a practical route to a more reliable response.

If speed is the priority, Terra 5.6 produces a shorter, accurate response in 40 minutes. The trade-off is depth. Across the charts, the decision depends on the work the user needs completed: quality, cost and speed each reveal something the other scores cannot.

These findings come from a specific RFP and 1 graded run per model and setup. Repeated tests are needed to distinguish recurring behaviour from variation between runs. We will keep running and publishing evals across RFPs and other sales tasks, and update our choices as the evidence changes.

What this means for our users

Evals help us do more than choose a model. They show us where to improve the whole agent: how it reads sources, handles contradictions, respects the user's decisions and turns its analysis into finished documents. Each recurring failure becomes a concrete target for our skills, prompts and tools. The goal is to reduce the checking and rework our customers need to do, so they can spend more time winning deals.

FAQ

What is the best AI model for writing RFP responses?

In Cobl's September 2026 benchmark on a real public-sector RFP, Astra 6 scored highest on quality at 87/100, followed by Opus 5 at 79/100 and Fable 5.1 at 74/100. Cobl uses Sol 5.6 as its default model: it scored 73/100 with zero fabricated facts across the month's tests and mid-range cost and delivery time.

How long does it take AI to write a full RFP response?

Among the runs through Cobl, recorded delivery time ranged from 40 minutes with Terra 5.6 to 126 minutes with Fable 5.1. Astra 6's session window was 65 minutes.

What is an LLM eval?

An eval is a repeatable test. Several AI models receive the same RFP, knowledge base and scripted user messages, and their documents, conversations and traces are assessed against a consistent set of rules. The result is a quality ranking, a cost and speed report, and a list of problems assigned to the people responsible for fixing them.

What are the most common AI mistakes in RFP responses?

In Cobl's test, the recurring failures were facts or prices not supported by the approved sources, problems raised in chat but left out of the final document, a missing penalty table, and documents that were never exported. Each of these can disqualify a bid or leave extra work for the user.

How much does it cost to generate an RFP response with AI?

Among the runs through Cobl, the total cost of a full RFP response ranged from $0.38 with Luna 5.6 to $68.09 with Astra 6. Sol 5.6, Cobl's default model, cost $8.05. Luna's price is provisional.