PROJECT GUIDE / METHODOLOGY & CONTEXT
Behind the
benchmarks.
What the tests measure, where the numbers come from and how to interpret them.
Introduction
Framework Benchmarks compares web stacks performing common tasks: JSON serialization, database access, cached reads and server-rendered HTML. The project began in March 2013 and grows through community-contributed implementations. TechEmpower publishes result rounds; Better Web maintains this independent portal and fork.
Motivation
Performance influences hosting capacity, infrastructure cost and responsiveness. Measuring the foundational stack can reveal constraints before architecture becomes more complex. These synthetic tasks support investigation, while application-specific testing and development needs complete the decision.
01 / ENVIRONMENT
The hardware changed.
Read each round in context.
The original suite separates application server, database server and load generator. Physical and cloud environments have evolved independently.
ROUNDS 23 →Citrine / 40 GbE
Three HPE ProLiant DL360 Gen10 Plus servers; Xeon Gold 6330 at 2 GHz, 56 cores as described upstream, 64 GB RAM, enterprise SSD and ConnectX-6 40 Gbps networking. Provided by Microsoft.
ROUNDS 16–22Citrine / 10 GbE
Dell R440 servers with Xeon Gold 5120, 32 GB RAM, enterprise SSD and a dedicated Cisco 10 Gbps switch. Provided by Microsoft.
ROUNDS 13–15ServerCentral
Dell R910 application server with four 10-core Xeon E7-4850 processors; Dell R710 database server with two 4-core Xeon E5520 processors; 10 Gbps Ethernet.
ROUNDS 9–12Peak Hosting
Dell R720xd, dual Xeon E5-2660 v2 (40 hyperthreads), 32 GB RAM, SSD RAID for databases and 10 Gbps networking.
ROUNDS 1–8Office i7
Core i7-2600K workstations, 8 GB RAM and gigabit Ethernet; Samsung 840 Pro database SSD from round 5.
ROUNDS 13 →Azure
Azure D3v2 instances and gigabit Ethernet, as documented in the upstream environment guide. Availability varies by round.
ROUNDS 1–12AWS
EC2 c3.large with 2 vCPUs; m1.large through round 9. Gigabit Ethernet.
Original environment specifications ↗02 / THE PROCEDURE
Configure. Verify. Warm up. Measure.
- Restart databases and start the chosen application.
- Run a short primer to confirm that the server responds.
- Warm up initialization and JIT compilation before measurement.
- Measure at the published concurrency levels or query counts using wrk.
- Capture throughput, latency, errors and validation outcomes.
- Stop the application before moving to the next implementation.
Round 23 data contains concurrency levels 16, 32, 64, 128, 256 and 512; plaintext levels 256, 1,024, 4,096 and 16,384; query counts 1, 5, 10, 15 and 20. Historical levels and duration are read directly from each snapshot.
The seven workload requirements ↗03 / TERMINOLOGY
A shared vocabulary.
- framework
- An HTTP stack used to build web applications, from a platform to a micro or full-stack framework.
- platform
- The runtime or server layer below a framework: for example Netty, Servlet or Rack.
- permutation
- One combination of language, framework, runtime, database and configuration.
- test type
- A workload such as JSON, database access or template rendering.
- test
- A measurement of one implementation executing one workload.
- implementation
- The application code and configuration that fulfill the workload requirements.
- toolset
- The Python runner and utilities that prepare, verify and measure implementations.
- run
- One execution of all or part of the benchmark suite.
- preview
- An intermediate result capture reviewed by project contributors.
- round
- A published TechEmpower result release. A round is distinct from a repository snapshot.
04 / QUESTIONS & INTERPRETATION
Before you choose a stack.
How should I choose a framework?
Use performance alongside language preference, team experience, maintainability, documentation and support. Validate the shortlist against your application workload and hardware.
Why compare platforms and full-stack frameworks?
They expose different abstraction costs. Filter by classification, database, ORM and platform when you need comparable stacks; the unfiltered table is deliberately broad.
What do Realistic and Stripped mean?
Realistic targets a general-purpose production configuration. Stripped is specially configured or built for benchmark requirements and can be unsuitable as a production starting point. Stripped is hidden initially, as in the original viewer.
What does Did not complete mean?
The implementation did not complete the workload successfully or pass the required validation. A failed test is shown as a failure, rather than assigned a performance score. Follow its logs to investigate.
Why are some tests or variants missing?
Workload support and included implementations change between rounds. A configuration in the current repository does not imply that it was measured in a historical publication.
Can different rounds be compared directly?
Hardware, network, software versions and implementations change over time. Treat each round and environment as a separate experiment. This portal does not combine them into a single score.
How are throughput and latency presented?
Published throughput uses the round duration and completed request counts after connection, read, write and HTTP 5xx errors. Latency comes from the same selected sample. For query workloads, select the query count; for concurrency workloads, inspect all levels as well as the best sample.
Why are local imports marked approximate?
The local results tool divides totalRequests by measured elapsed timestamps. The published-round viewer uses the published duration and error-adjusted counts. These are different calculations, so their values can differ.
What is framework overhead?
It compares an implementation with its explicitly declared platform baseline in the same test and environment. The percentage here is framework RPS divided by baseline RPS. No baseline is inferred from a similar name or language.
How does the runner handle warmup?
The upstream procedure describes a 5-second primer at concurrency 8, a 15-second warmup at 256 and measured samples at the workload levels. Primer and warmup are excluded from published measurements. Check runner settings when reproducing a run.
Are 15-second samples representative?
The suite balances measurement time against a large number of implementations. A short synthetic run cannot represent every long-lived application, GC pattern or JIT behavior. Extend duration and repeat runs for your own investigation.
What does wrk do?
wrk is the multithreaded HTTP load generator. It issues requests for a configured duration and records request counts, latency and errors. Load-generator or network saturation can limit the observed application throughput.
Are caches and reverse proxies used?
Cached queries is a separate workload with its own requirements. Ordinary database tests must perform their specified database operations. Responses served without executing application work by an external reverse cache would change what is being measured.
What about pooling, logging and CPU usage?
The upstream expectations include connection pooling and configurations that use the available CPU resources. Application logging is generally disabled in the benchmark. Review each implementation rather than assuming all production defaults are enabled.
Why are versions or configurations outdated?
The suite evolves through community contributions. Report configuration problems with a reproducible example or submit a change with verification results. A published round remains a historical snapshot even after the repository is updated.
How can I add a framework or test?
Start from the workload requirements, contribute an implementation and verify it with the runner. Use the fork for Better Web changes and the upstream project for changes intended for official TechEmpower rounds.
Where can contributors see continuous runs?
TFB Status publishes continuous benchmark runs and logs. These are useful for diagnosing contributions and are distinct from official published rounds.
Where are rounds 1 and 2?
The original viewer starts its historical data at round 3 because early result formats and identifiers changed. The first two releases remain available as TechEmpower blog posts below.
05 / PREVIOUS ROUNDS
The publication archive.
Open the available published data inside the new viewer. Round 12 and 14 cloud data were unavailable at the upstream endpoint when this snapshot was collected; the original publication remains linked.
Round 1 · March 2013 ↗Round 2 · April 2013 ↗
06 / SOURCE & COMMUNITY
Inspect it. Improve it.
This guide paraphrases and organizes the original TechEmpower project information. Published data and historical hardware facts are attributed to TechEmpower. Refer to the original documentation for complete, current requirements. Original site ↗