1 Definition and purpose
Proficiency testing is a structured process used to assess a laboratory’s performance by comparing its results with those of other participants or with assigned reference values. It serves as an external check on analytical quality and helps determine whether a laboratory can produce reliable results under routine conditions. In many fields, it is one of the clearest ways to demonstrate technical competence.
1.1 Core concept
The core idea is simple: multiple laboratories examine the same or equivalent test material, then their findings are compared. If a laboratory’s result falls within an acceptable range, it indicates that the method and procedures are functioning as expected. If the result differs substantially, the discrepancy may point to an error in calibration, technique, calculation, or sample handling.
1.2 Distinction from internal quality control
Internal quality control takes place within a laboratory and monitors ongoing routine performance, often by using control materials, charts, and repeat measurements. Proficiency testing is external and comparative. It does not merely show whether an instrument is stable from day to day; it shows how a laboratory performs relative to peers or a reference target. For that reason, the two approaches complement each other.
1.3 Role in quality assurance
Within a broader quality assurance system, proficiency testing provides independent evidence that a laboratory’s results are trustworthy. It can reveal problems that routine internal checks may miss, especially those involving operator bias, hidden method flaws, or mistaken assumptions about a procedure. The findings are often used in audits, inspections, and accreditation reviews.
1.4 Typical objectives
Common objectives include verifying competence, detecting analytical bias, comparing methods, and supporting continual improvement. Programs may also be designed to evaluate specific analytes, challenge difficult matrices, or confirm that a laboratory can meet regulatory or contractual requirements. In practice, proficiency testing is both a measurement tool and a management tool.
2 Historical development
Proficiency testing grew from the long-standing scientific practice of comparing observations between laboratories. As analytical methods became more standardized, formal schemes were created to reduce variability and strengthen confidence in results. Over time, these comparisons became integral to laboratory governance and accreditation.
2.1 Early interlaboratory comparisons
Early interlaboratory comparisons were often informal and focused on establishing whether different laboratories could reproduce similar results. Such exercises were especially valuable in chemistry, medicine, and measurement science, where precision and reproducibility were essential. These comparisons helped expose inconsistencies in methods and encouraged standardization.
2.2 Adoption in standards and accreditation
As laboratory quality systems matured, proficiency testing was incorporated into standards, guidance documents, and accreditation frameworks. It became a recognized means of demonstrating technical competence, particularly in fields where results affect public health, product safety, or legal compliance. Formal schemes also introduced scoring rules and performance criteria, making evaluation more systematic.
2.3 Expansion into modern laboratory systems
With the growth of specialized testing fields and international commerce, proficiency testing expanded into many sectors. Today it is used by clinical, environmental, food, industrial, and research laboratories across different countries and disciplines. Digital reporting, statistical software, and global provider networks have made large-scale programs easier to administer.
3 Proficiency testing schemes
A proficiency testing scheme is the organized framework that defines how samples are prepared, distributed, analyzed, and evaluated. The scheme must be designed so that the challenge is fair, the materials are suitable, and the results can be interpreted consistently. Its credibility depends on careful administration and transparent rules.
3.1 Organizing bodies
Schemes may be run by professional organizations, accreditation-related groups, government agencies, reference laboratories, or commercial providers. The organizer is responsible for planning the exercise, preparing instructions, assigning values where appropriate, and issuing reports. Reliable organization is essential because participants depend on the scheme’s neutrality and technical rigor.
3.2 Participant enrollment
Enrollment usually requires a laboratory to register for one or more analytes, methods, or testing disciplines. Participants may join voluntarily, or they may be required to take part because of accreditation, regulation, or contractual obligations. Clear enrollment terms help ensure that all laboratories understand the scope, timeline, and expectations of the exercise.
3.3 Sample selection and preparation
Test materials must resemble the type of specimens or products a laboratory normally handles. They may be natural samples, fortified materials, or specially prepared matrices designed to test a specific measurement. Good preparation aims for homogeneity, stability, and enough similarity to routine samples that the exercise reflects real-world performance.
3.4 Distribution and handling of test materials
Once prepared, the materials are packaged and sent to participants under conditions that protect their integrity. Instructions typically specify storage, handling, analysis windows, and reporting deadlines. If a material is sensitive to temperature, time, or contamination, the organizer must control these factors carefully to preserve the fairness of the exercise.
4 Testing process
The testing process follows a sequence that allows laboratories to analyze the same challenge material independently and then submit their findings for comparison. Although the details vary by field, the overall structure is similar across most proficiency testing programs. The process is designed to mirror ordinary laboratory practice.
4.1 Analysis by participating laboratories
Participants analyze the samples using their routine procedures, instruments, and personnel rather than special research methods. This approach helps reveal how the laboratory truly performs in normal operation. Laboratories are usually expected to treat the proficiency test material as they would a regular client sample.
4.2 Result submission
After analysis, each laboratory submits its results in the required format, often along with method details, units, and supporting information. Accurate transcription and unit consistency are important because reporting errors can distort the evaluation. Some programs also ask for qualitative observations, uncertainty estimates, or detection limits.
4.3 Statistical evaluation
The organizer compiles the submitted results and evaluates them using predetermined statistical procedures. Depending on the design, the comparison may be against a reference value, a consensus value, or a set distribution of participant results. Statistical treatment helps identify both typical performance and unusual deviations.
4.4 Performance reporting
Participants receive a report showing their result, the assigned value or comparison point, and the performance assessment. Reports may include scores, graphs, method summaries, and comments on trends over time. Many schemes provide confidential individual feedback while also offering a general summary of overall participation.
5 Evaluation methods
Evaluation methods are used to determine whether a laboratory’s result is acceptable and how far it differs from the target value. The choice of method depends on the analyte, the purpose of the scheme, and the statistical properties of the dataset. Good evaluation methods balance fairness, clarity, and sensitivity to error.
5.1 Assigned values
An assigned value is the reference point against which participant results are compared. It may be derived from a certified reference material, a reference method, a consensus of expert laboratories, or a carefully characterized target value. The reliability of the assigned value strongly affects the credibility of the assessment.
5.2 Statistical measures
Common statistical measures include mean, median, standard deviation, robust estimates, and deviation from the assigned value. These measures help summarize group performance and identify outliers. In some schemes, uncertainty and measurement dispersion are also considered when judging acceptability.
5.3 Z-scores and related indicators
A widely used indicator is the z-score, which expresses how far a result lies from the assigned value in units of standard deviation. Related indicators may adjust for uncertainty, bias, or method-specific criteria. These scores provide a compact way to classify performance as satisfactory, questionable, or unsatisfactory.
5.4 Acceptance criteria
Acceptance criteria define the threshold at which a result is considered acceptable. They may be based on biological relevance, technical capability, regulatory needs, or fitness for purpose. Clear criteria are important because a result that is statistically unusual is not always practically unacceptable, and vice versa.
6 Applications
Proficiency testing is used wherever measurement quality matters and where comparisons can strengthen confidence in results. Its role differs by sector, but the underlying purpose remains the same: to demonstrate that a laboratory can produce dependable data. The breadth of applications has made it a central feature of modern laboratory practice.
6.1 Clinical laboratories
In clinical laboratories, proficiency testing helps verify the accuracy of tests used in diagnosis, treatment monitoring, and screening. It can cover chemistry, hematology, microbiology, immunology, and other disciplines. Because patient care depends on trustworthy results, this is one of the most established uses of the practice.
6.2 Environmental testing
Environmental laboratories use proficiency testing to assess measurements of water, air, soil, and waste samples. These programs help confirm that the laboratory can detect contaminants and quantify pollutants accurately. Reliable results are essential for monitoring, permitting, and compliance work.
6.3 Food and agricultural analysis
Food and agricultural laboratories rely on proficiency testing for nutrient analysis, contamination detection, residue testing, and authenticity checks. The matrices in these programs can be complex, which makes external comparison particularly useful. Good performance supports consumer protection and product consistency.
6.4 Industrial and manufacturing laboratories
Industrial laboratories use proficiency testing to confirm the quality of materials, components, and process measurements. This may include metal analysis, chemical composition, product specification testing, and process control verification. Comparable results help manufacturers maintain consistency across sites and suppliers.
6.5 Research and metrology
In research settings, proficiency testing can support method validation and collaborative projects. In metrology, it is especially important for comparing measurement capability and traceability. These applications often involve highly controlled materials and strict statistical treatment.
7 Standards and accreditation
Standards and accreditation frameworks frequently require or recommend proficiency testing as evidence of competence. These systems use external comparison to support confidence in laboratory outputs and to encourage continuous improvement. The exact requirements vary by discipline and governing body.
7.1 International standards
International standards often describe how proficiency testing should be organized, how results should be evaluated, and how participants should respond to poor performance. Such standards promote consistency across schemes and make results more comparable across borders. They also help define the responsibilities of providers and users.
7.2 Regulatory requirements
In regulated fields, proficiency testing may be mandated for certain tests, laboratories, or sectors. Regulatory rules can specify participation frequency, acceptable scoring systems, and reporting obligations. These requirements are intended to protect public interests where inaccurate data could have serious consequences.
7.3 Accreditation body expectations
Accreditation bodies often expect laboratories to participate in relevant proficiency testing programs as part of their quality system. They may review participation records, investigate failures, and examine how corrective actions were implemented. A pattern of poor performance can affect accreditation status or trigger closer oversight.
7.4 Documented competence
Documented competence means that a laboratory can show evidence of maintaining technical capability over time. Proficiency testing provides one form of that evidence because it records performance against external benchmarks. When combined with training records, method validation, and internal controls, it strengthens the overall competence profile.
8 Challenges and limitations
Although proficiency testing is valuable, it has practical and statistical limitations. Not every test material behaves like a real sample, and not every method is directly comparable. Interpreting results requires attention to context and to the design of the scheme.
8.1 Matrix effects
Matrix effects occur when the sample’s composition influences the measurement result. A material that is easy to analyze in one matrix may be difficult in another, leading to performance differences that reflect the sample rather than the laboratory alone. Good schemes try to match the matrix as closely as possible to routine specimens.
8.2 Sample homogeneity and stability
For a scheme to be fair, each participant must receive essentially the same material, and that material must remain unchanged during transport and storage. If a sample is not homogeneous or degrades over time, the comparison becomes less reliable. Organizers therefore test these properties before distribution.
8.3 Comparability of methods
Different analytical methods can produce systematically different results even when all are performed correctly. This makes it difficult to judge performance using a single universal target in some fields. Method-specific evaluation or grouped comparison may be needed to avoid misleading conclusions.
8.4 Interpretation of poor performance
A poor score does not automatically reveal the source of the problem. The issue could stem from calibration, contamination, transcription, instrument drift, sample preparation, or an unsuitable method. Careful interpretation is necessary so that the result leads to useful corrective action rather than a superficial conclusion.
9 Corrective and preventive action
When a laboratory performs poorly in proficiency testing, the result should trigger a structured response. The goal is not only to fix the immediate issue but also to reduce the chance of recurrence. This is an important part of laboratory quality management.
9.1 Investigation of deviations
The first step is to examine the deviation in detail and confirm that it is real. The laboratory may review the original records, check calculations, and compare the result with internal control data. This investigation helps determine whether the discrepancy was isolated or part of a broader pattern.
9.2 Root-cause analysis
Root-cause analysis seeks the underlying reason for the failure rather than the visible symptom. Common causes include incorrect calibration, reagent problems, operator error, poor sample preparation, or unsuitable method settings. Identifying the cause makes the response more effective and prevents repeated mistakes.
9.3 Method review and retraining
If the investigation suggests a procedural weakness, the laboratory may revise the method, adjust verification steps, or retrain staff. Instrument maintenance and documentation practices may also be updated. These actions strengthen the laboratory’s routine workflow and reduce future risk.
9.4 Follow-up testing
After corrective measures are implemented, follow-up testing can confirm whether performance has improved. This may involve the next proficiency round, a targeted comparison, or internal rechecking. Follow-up is important because it closes the loop between diagnosis and resolution.
10 Related concepts
Proficiency testing is closely related to several other forms of comparison and quality assessment. These concepts overlap in purpose but differ in design, scope, and formal use. Understanding the distinctions helps clarify how laboratories evaluate performance.
10.1 Interlaboratory studies
Interlaboratory studies are broader comparisons among laboratories, often used to examine methods, precision, or reproducibility. They may be research-oriented rather than evaluative. Proficiency testing can be considered a more structured and performance-focused form of interlaboratory comparison.
10.2 Round-robin testing
Round-robin testing is a coordinated exercise in which a sample or task is circulated among participants in sequence. It is commonly used to compare methods or assess variability across laboratories. The term is often used interchangeably with interlaboratory comparison, though the exact usage may vary.
10.3 External quality assessment
External quality assessment is a wider category that includes proficiency testing and similar external reviews of laboratory performance. In some contexts, the two terms are used nearly synonymously. In others, external quality assessment may refer to a broader program that includes more than a single comparison exercise.
10.4 Reference materials
Reference materials are well-characterized substances used to calibrate instruments, validate methods, or check analytical performance. They may be certified for one or more properties and are often central to assigned values in testing schemes. Their use supports traceability and consistency across laboratories.