GPT-4o Performance: August vs. November Versions 2024 — Unique AI Documentation

GPT-4o Performance: August vs. November Versions 2024

Purpose

This benchmarking report evaluates two distinct versions of OpenAI’s GPT-4o model:

The objective is to assess the relative quality and consistency of responses as the model evolves. To eliminate potential evaluator bias, all evaluations were conducted as blind tests, where the model version was not disclosed to reviewers.


Performance Overview – Model Consistency Check

This section evaluates the internal consistency of each model by comparing responses between two separate runs using the same 141 prompts.

The comparison is conducted using our LLM-based benchmarking tool, where an LLM acts as a judge, scoring outputs across four critical dimensions:

Consistency Metrics Across Model Versions

Metric GPT-4o-08 GPT-4o-11
Contradiction 16% 15%
Extent Difference 18% 14%
Hallucinations 4% 4%
Source Variations 32% 39%

GPT-4o-11 shows a slightly better consistency across contradiction and extent difference dimensions, while both models perform equally well regarding hallucinations. However, GPT-4o-11 introduces more source variation, suggesting differences in source citation or retrieval alignment.


Qualitative Evaluation – Cross-Model and Intra-Model Analysis

This evaluation involves human reviewers comparing responses from:

  1. GPT-4o-08 vs GPT-4o-11 (cross-version comparison)
  2. GPT-4o-11 vs GPT-4o-11 (intra-version consistency)

Each model was run twice per question, and reviewers compared the two generated outputs for differences in meaning, style, and alignment to sources.

Evaluation Metrics

Comparison Type Obvious Differences, Same Meaning Very Slight Differences Identical Responses Meaningfully Different
GPT-4o-08 vs GPT-4o-11 34% 57% 6% 2%
GPT-4o-11 vs Itself 26% 55% 18% 1%

Insights


Behavioral Patterns

Behavioral distinctions emerged between the two model versions, especially in tone, structure, and contextualization:

Examples

Q: Have there been any adverse ESG events communicated to investors?

Q: Is the valuation policy board-approved?


Conclusion

GPT-4o-11 (November 2024 release) provides more reliable, structured, and source-aligned responses. It consistently delivers answers that match institutional standards for clarity and accuracy, especially in compliance-driven or complex domains.

While GPT-4o-08 (August 2024 release) offers faster readability and efficiency — making it valuable for routine tasks and short-form outputs — GPT-4o-11 is the preferred choice when accuracy, justification, and regulatory alignment matter.


Recommendation

Based on this evaluation:

Users should also run independent benchmarks to verify model performance in their unique use cases.


Retrieval-Augmented Generation (RAG) Configuration

The RAG configuration was identical for both model versions, ensuring a fair and consistent comparison environment.

Key Parameters