Go back to all case studies

A Platform Specified for 500,000 Concurrent Users Reached 2% of It

Case Study
7 min read
Denis Sautin

Denis Sautin

Author

Denis Sautin

Denis Sautin is an experienced Product Marketing Specialist at PFLB. He focuses on understanding customer needs to ensure PFLB’s offerings resonate with you. Denis closely collaborates with product, engineering, and sales teams to provide you with the best experience through content, our solutions, and your personal journey on our website.

Product Marketing Specialist

Reviewed by Boris Seleznev

boris author

Reviewed by

Boris Seleznev

Boris Seleznev is a seasoned performance engineer with over 10 years of experience in the field. Throughout his career, he has successfully delivered more than 200 load testing projects, both as an engineer and in managerial roles. Currently, Boris serves as the Professional Services Director at PFLB, where he leads a team of 150 skilled performance engineers.

At a glance

  • Client: a company delivering a national education project built on the platform-as-a-service model - a platform for interaction between teachers and students
  • Their deadline: a pilot in general education schools across six regions of the country in six months, then replication to general education schools nationally
  • How the platform was sourced: the customer chose to buy and customise rather than build in house, purchasing a solution from an American educational platform developer and enhancing it to meet the requirements
  • The requirement under test: maximum performance supporting as many as 500,000 simultaneous users
  • What the load stations showed first: CPU utilisation on the load station never exceeded 20%, then 25% after a move to a gigabit network - the equivalent of 3% of target load
  • What the stand showed once monitoring was installed: at 10% of target load the web server's CPU reached 100%, dropping to 2% the moment the test stopped
  • Methodology change forced mid-project: the planned 10%-increment stages (10% → 110%) were replaced by twenty 0.5% micro-stages up to 10%
  • Result: measured maximum performance was 2% of the target - with the causes localised, described, and delivered early enough for the customer to still change vendor

Six regions, six months, and no users yet

In the autumn a company handed PFLB a load testing task. It turned out that the customer was building an educational project on the platform-as-a-service principle - a platform meant to make interaction between teachers and students work properly, at national scale.

The commitment behind it was the part that mattered. A pilot had to run in general education schools across six regions of the country within six months, and on the strength of that pilot the platform would be replicated to general education schools generally. There were two ways to get there: buy a box solution and customise, refine and integrate it, or build an in-house information system. The customer chose the first, purchasing a solution for the pilot from an American educational platform developer and enhancing the base product to meet the stated requirements.

Which left one open question, and it was a big one: the requirement said the system had to support as many as 500,000 simultaneous users. Nobody had confirmed that it did. PFLB proposed starting with a single iteration of load testing to determine maximum performance - not a full programme, one measurement of the ceiling.

Five things that had to happen before any load was generated

A maximum performance test is not something you start on day one. Five preparatory tasks came first:

1. Describe the system architecture. 2. Define the characteristics of the testing stand and assemble it. 3. Define the list of business processes that go into the load testing profile. 4. Develop the load testing profile from the data actually available. 5. Write the load testing methodology, specifying the mechanism of the test in detail.

The stand was the first real judgement call. The ideal configuration is an identical copy of the industrial stand, which is rarely achievable for a large-scale system; the job is to pick a configuration that minimises the capacity consumed while still producing data that extrapolates reliably to production. Here PFLB engineers were lucky - the system was not yet in industrial operation, which gave a rare chance to test in a prepared "battlefield environment".

The same fact was also the project's biggest handicap. With no industrial operation there were no user statistics to build a profile from. Where a profile is compiled expertly rather than measured, the business customer and industry experts are normally pulled in; here the team also had access to open sources and used official data from the Ministry of Education and Science. Out of that came a finished load profile, defined requirements for filling the database, and a prepared stand.

Measuring the load generators before measuring the system

Methodology writing and the preparation of load testing facilities ran in parallel from the start of the project. Load testing facilities are the capacity that generates the load, and before they can be trusted the throughput of a single load station has to be established - that number is what determines how many stations are needed to deliver the target load.

Load station throughput is established through synthetic testing: standardised tests that report the performance of an IT system in hardware and software terms.

At this stage the team had no access to the testing stand machines and could not observe the stand's own hardware utilisation. The results were strange:

  • After the first synthetic test, CPU utilisation at the load station did not exceed 20%. The obvious suspect was non-optimal load scripts.
  • A series of script optimisations followed. No improvement.
  • The next hypothesis was the network channel - the scripts carried a large volume of web statics. The load stations were moved onto a single network with a gigabit channel instead of 100 Mbps.
  • Several more synthetic tests returned a similar picture: load station CPU under 25%, equal to 3% of the target load.

With every cause on the load generation side eliminated, access to the testing stand machines was granted in due time and monitoring software for hardware metrics was installed on the stand itself. Synthetic testing continued - and the answer appeared immediately.

What the stand was actually doing

At 10% of the target load, the web server's CPU utilisation reached 100%. When the test was terminated, that same CPU utilisation dropped to 2%.

The customer was notified at once, and the test plan and load testing methodology had to be corrected quickly. The methodology as written searched for maximum performance in stages of 10% of target intensity, from 10% up to 110% - a sensible design for a system that gets somewhere near its requirement, and a useless one here, because the first stage already saturated the web server. Starting at 10% no longer made sense.

The fix was to introduce "micro stages" of 0.5% of the target load, twenty of them, covering the range up to 10% - enough resolution to characterise a system that failed inside the first step of the original plan.

Results

  • Maximum system performance measured at 2% of the target performance. The configuration as delivered was nowhere near the stated requirement of 500,000 simultaneous users.
  • The causes of the performance limitation were localised, and recommendations for eliminating them were described in detail.
  • The customer obtained the picture early and cheaply - the limits of the system, the located bottlenecks, and a realistic estimate of the work required to raise performance.
  • The finding arrived while the decisions were still open. Because testing happened at an early stage, replacing the vendor or developer remained an acceptable option, and PFLB was positioned to support the removal of the constraints and repeat the test iteration.

Questions this engagement answers

How long does a load testing engagement take?

It is governed by preparation and by what the first measurements find, not by the test itself. Here five preparatory tasks - architecture description, stand assembly, business process selection, profile development and methodology - came before any load was generated, and the schedule then changed again mid-project when the first synthetic tests showed the web server saturating at a tenth of target and the 10%-step methodology had to be rewritten as twenty 0.5% micro-stages.

Can you build a load profile for a system that has no users?

Yes, but not from logs. With the system outside industrial operation there were no usage statistics at all, so the profile was assembled from expert input and official open data from the Ministry of Education and Science.

Why test the load generators before testing the system?

Because you cannot tell a slow system from a slow rig without doing it. Two rounds of investigation here - script optimisation, then a move from 100 Mbps to gigabit - were needed to prove the load stations were not the limitation before the stand could be blamed.

Is it worth testing before the platform goes live?

This case is the argument. A system delivering 2% of its required capacity is a commercial problem, not a tuning problem - and it was found while changing vendor was still on the table rather than after a six-region pilot had started.

Buying a platform against a performance requirement?

If a contract, a pilot deadline or a national rollout depends on a stated capacity figure, the useful moment to measure it is before the rollout starts - while the vendor decision, the architecture and the schedule are all still open.