EFQM is an assessment framework, not a theory, and saying so at the start is not pedantry. A theory makes a claim that could be false. EFQM makes no such claim. It supplies a criterion set, a set of weights, a procedure for gathering evidence and a way of turning that evidence into a number, and its authority comes from adoption rather than from evidence.
What it actually contains is an architecture with four parts: a published list of criteria describing how a well-run organization behaves and what it achieves; fixed point weights that say how much each criterion counts; an assessment logic called RADAR that tells an assessor how to convert observed practice into a percentage; and a corps of trained assessors who apply it to a written application, usually with a site visit. The output is a score out of one thousand and a report naming strengths and areas for improvement.
Its usefulness is real and it is organizational rather than scientific. It gives a management team a shared vocabulary, forces the writing down of approaches that were tacit, and produces a comparison against a standard that is external to the organization. Those are genuine goods. None of them requires the model to be true, because there is nothing in it that could be true or false in the way a theory is.
For a research student the practical consequence is a warning. EFQM will not carry a thesis as its theoretical framework. It can serve as an instrument to be validated, as an intervention whose effect is estimated, or as a source of a dataset of assessed organizations. Those are three good uses. Presenting it as the theory the study rests on is the mistake the literature makes most often.
نموذج EFQM إطارُ تقويمٍ لا نظرية، وقولُ ذلك ابتداءً ليس تشدّقًا. فالنظرية تقول دعوى يمكن أن تكون كاذبة. وهذا النموذج لا يقول دعوى من هذا النوع. بل يقدّم مجموعةَ معايير، ومجموعةَ أوزان، وإجراءً لجمع الأدلّة، وطريقةً لتحويل الأدلّة إلى رقم، وسلطتُه آتيةٌ من الانتشار لا من البيّنة.
والذي يحويه فعلًا بناءٌ من أربعة أجزاء: قائمةٌ منشورة من المعايير تصف كيف تسلك المنظمة حسنةُ الإدارة وماذا تحقّق؛ وأوزانٌ ثابتة من النقاط تقول كم يزن كل معيار؛ ومنطقُ تقويمٍ يسمّى رادار يخبر المقيِّم كيف يحوّل الممارسة الملاحَظة إلى نسبة مئوية؛ وسلكٌ من المقيِّمين المدرَّبين يطبّقونه على طلبٍ مكتوب، مع زيارةٍ ميدانية في العادة. والمخرَجُ درجةٌ من ألف، وتقريرٌ يسمّي مواطن القوّة ومجالات التحسين.
ونفعُه حقيقيّ، وهو نفعٌ تنظيمي لا علميّ. فهو يعطي فريق الإدارة مفرداتٍ مشتركة، ويُلجئ إلى كتابة مناهجَ كانت ضمنية، وينتج مقارنةً بمعيارٍ خارجٍ عن المنظمة. وهذه خيراتٌ صادقة. ولا يقتضي شيءٌ منها أن يكون النموذج صادقًا، إذ ليس فيه ما يمكن أن يصدق أو يكذب كما تصدق النظرية أو تكذب.
واللازم العملي لطالب البحث تحذير. فـ EFQM لن يحمل رسالةً بوصفه إطارها النظري. وهو يصلح أداةً تُصدَّق، أو تدخّلًا يُقدَّر أثرُه، أو مصدرًا لبياناتٍ عن منظماتٍ مقوَّمة. وهذه ثلاثة استعمالاتٍ حسنة. وعرضُه نظريةً تقوم عليها الدراسة هو الخطأ الذي تقع فيه الأدبيات أكثر من غيره.
The Deming Prize, 1951. The oldest of the award-based quality assessments, established in Japan and named for a statistician whose lectures on process control were taken more seriously there than at home. It set the pattern that everything after it followed: an application document, an external examination, and public recognition tied to management practice rather than to a product test.
Baldrige, 1987. The Malcolm Baldrige National Quality Award in the United States built the architecture EFQM inherited almost intact. A published criterion set with fixed point weights, a written application, trained volunteer examiners, a site visit, a score out of one thousand, and a national award presented at the top of government. The design choice that mattered most was making the criteria public, which turned the award into a self-assessment instrument used by thousands of organizations that never applied.
The founding, 1988 to 1991. Fourteen European companies established the European Foundation for Quality Management in 1988. The model and the European Quality Award followed in 1991, first awarded in 1992. The motivation was competitive rather than scholarly: European manufacturers wanted an answer to Japanese quality practice and a European counterpart to the American award.
The 1999 revision and the revisions after it. The 1999 version renamed the framework the EFQM Excellence Model, introduced RADAR as the assessment logic in place of the earlier scoring matrix, and settled the nine criteria into five enablers and four results weighted five hundred points against five hundred. Further revisions followed in 2003, 2010 and 2013, adjusting criterion parts, the fundamental concepts and the weights. None of the reweightings was published alongside an estimation that would justify it.
The 2020 restructure. The nine-box picture was abandoned. The model was rebuilt around three blocks, Direction, Execution and Results, holding seven criteria, with a stronger emphasis on purpose, transformation and ecosystem thinking. RADAR was retained and rewritten. The change was substantial enough that scores under the new structure are not comparable with scores under the old one, which is rarely stated when trend charts are drawn.
Bou-Llusar and colleagues, 2009. The most serious empirical examination, testing the enabler and results structure with structural equation models on assessment data and comparing the model against the Baldrige criteria. The support was partial. The enabler constructs were highly collinear, the assumed separation between criteria did not reproduce cleanly, and the fit of the postulated causal ordering was weaker than its universal use in practice implies.
جائزة ديمنغ، ١٩٥١. أقدمُ تقويمات الجودة القائمة على الجوائز، أُنشئت في اليابان وسُمّيت باسم إحصائيٍّ أُخذت محاضراتُه في ضبط العمليات هناك مأخذًا أجدَّ مما أُخذت في بلده. وقد أرست النمط الذي تبعه كل ما جاء بعدها: وثيقةُ طلب، وفحصٌ خارجي، وتكريمٌ علنيّ معلَّق بممارسة الإدارة لا باختبار منتج.
بالدريج، ١٩٨٧. بنت جائزةُ مالكوم بالدريج الوطنية للجودة في الولايات المتحدة البناءَ الذي ورثه EFQM كاملًا تقريبًا. مجموعةُ معايير منشورة بأوزانٍ ثابتة من النقاط، وطلبٌ مكتوب، وفاحصون متطوّعون مدرَّبون، وزيارةٌ ميدانية، ودرجةٌ من ألف، وجائزةٌ وطنية تُسلَّم في أعلى الدولة. وأهمّ اختيارٍ تصميمي فيها نشرُ المعايير، وهو ما حوّل الجائزة إلى أداة تقويمٍ ذاتي تستعملها آلافُ المنظمات التي لم تتقدّم إليها قطّ.
التأسيس، ١٩٨٨ إلى ١٩٩١. أنشأت أربع عشرة شركة أوروبية المؤسسةَ الأوروبية لإدارة الجودة سنة ١٩٨٨. وتبعها النموذجُ والجائزةُ الأوروبية للجودة سنة ١٩٩١، ومُنحت أول مرّة سنة ١٩٩٢. وكان الباعث تنافسيًّا لا علميًّا: أراد الصنّاع الأوروبيون جوابًا عن الممارسة اليابانية في الجودة، ونظيرًا أوروبيًّا للجائزة الأمريكية.
مراجعة ١٩٩٩ وما بعدها. سمّت نسخةُ ١٩٩٩ الإطارَ نموذج التميز EFQM، وأدخلت رادار منطقًا للتقويم بدل مصفوفة التسجيل السابقة، واستقرّت بالمعايير التسعة على خمسة ممكِّنات وأربع نتائج بوزن خمسمئة نقطة في مقابل خمسمئة. وتلتها مراجعاتٌ سنة ٢٠٠٣ و٢٠١٠ و٢٠١٣، عدّلت أجزاء المعايير والمفاهيم الأساسية والأوزان. ولم تُنشر واحدةٌ من إعادات الترجيح مقرونةً بتقديرٍ يسوّغها.
إعادة بناء ٢٠٢٠. هُجرت صورةُ الصناديق التسعة. وأُعيد بناء النموذج حول ثلاث كتل: التوجّه والتنفيذ والنتائج، تحوي سبعة معايير، مع تشديدٍ أقوى على الغاية والتحوّل والتفكير المنظومي. واحتُفظ برادار بعد إعادة صياغته. وكان التغيير من الجوهرية بحيث لا تقارَن درجاتُ البِنية الجديدة بدرجات القديمة، وقلّما يُذكر ذلك حين تُرسم منحنياتُ الاتجاه.
بو لوسار وزملاؤه، ٢٠٠٩. أجدُّ فحصٍ تجريبي، اختبر بِنية الممكِّنات والنتائج بنماذج المعادلات البِنيوية على بيانات تقويمٍ فعلية، وقارن النموذج بمعايير بالدريج. وجاء السند جزئيًّا. فبِناءات الممكِّنات كانت شديدة التداخل الخطّي، والفصلُ المفترَض بين المعايير لم يظهر نظيفًا، وملاءمةُ الترتيب السببي المفترَض كانت أضعف مما يوحي به استعمالُه الشامل في الممارسة.
The architecture. Four parts working together: criteria that describe practice and achievement, weights that price each criterion, an assessment logic that converts observation into a percentage, and assessors trained to apply that logic consistently. Remove any one and the score stops meaning anything. Most published discussion of EFQM addresses only the first part, which is why so much of it misses where the score actually comes from.
| Block | Criterion in the 2020 structure | Points |
|---|---|---|
| Direction | Purpose, vision and strategy | 100 |
| Direction | Organisational culture and leadership | 100 |
| Execution | Engaging stakeholders | 100 |
| Execution | Creating sustainable value | 200 |
| Execution | Driving performance and transformation | 100 |
| Results | Stakeholder perceptions | 200 |
| Results | Strategic and operational performance | 200 |
The older arrangement is still the one most literature was written about. From 1999 the model held nine criteria split into five enablers and four results, weighted five hundred points each. Among the enablers, leadership carried 100, policy and strategy 80, people 90, partnerships and resources 90, and processes 140. Among the results, customer results carried 200, key performance results 150, people results 90 and society results 60. The enablers were said to drive the results and the results were said to feed learning back into the enablers. Anyone reading an empirical paper on EFQM published before roughly 2020 is reading about this structure, not the current one.
RADAR is the part that does the work. The letters stand for Results, Approach, Deploy, Assess and Refine. For an enabler criterion the assessor scores whether the approach is sound and integrated with the strategy, whether it is deployed systematically across the relevant parts of the organization, and whether it is measured, learned from and improved. For a results criterion the assessor scores relevance and usability, meaning scope, segmentation and integrity of the data, and performance, meaning trends over several years, achievement against targets, comparison with relevant external benchmarks, and confidence that the level will be sustained.
Each RADAR element is scored in percentage bands from zero to one hundred, the elements are combined into one percentage per criterion, and that percentage is multiplied by the criterion's point weight. Recognition tiers are then defined by bands of the total. The arithmetic is simple, which is part of the appeal, and it hides the fact that every input to it is a trained human judgement about a written document.
Relation to Baldrige and to national awards. Baldrige and EFQM share their architecture and differ in their criterion names, their weights and their vocabulary. Baldrige places roughly 450 of its 1,000 points on results against 400 in the 2020 EFQM structure, and organizes enablers into seven categories rather than five criteria. Most national excellence awards, in Europe and well beyond it, are derived from one or the other, usually with local criteria appended. That derivation is what makes the family worth studying: the same instrument, applied by different bodies, to different populations, with different appended criteria.
البناء. أربعةُ أجزاء تعمل معًا: معايير تصف الممارسة والإنجاز، وأوزانٌ تسعّر كل معيار، ومنطقُ تقويمٍ يحوّل الملاحظة إلى نسبة، ومقيِّمون مدرَّبون على تطبيق ذلك المنطق باتّساق. وارفع واحدًا منها تتوقّف الدرجة عن أن تعني شيئًا. وأكثرُ ما يُنشر في EFQM يعالج الجزء الأول وحده، ولهذا يفوت كثيرًا منه من أين تأتي الدرجة فعلًا.
| الكتلة | المعيار في بِنية ٢٠٢٠ | النقاط |
|---|---|---|
| التوجّه | الغاية والرؤية والاستراتيجية | ١٠٠ |
| التوجّه | الثقافة المؤسسية والقيادة | ١٠٠ |
| التنفيذ | إشراك أصحاب المصلحة | ١٠٠ |
| التنفيذ | خلق قيمة مستدامة | ٢٠٠ |
| التنفيذ | دفع الأداء والتحوّل | ١٠٠ |
| النتائج | تصوّرات أصحاب المصلحة | ٢٠٠ |
| النتائج | الأداء الاستراتيجي والتشغيلي | ٢٠٠ |
والترتيب الأقدم هو الذي كُتبت عنه أكثرُ الأدبيات. فمنذ ١٩٩٩ حمل النموذج تسعة معايير مقسومةً إلى خمسة ممكِّنات وأربع نتائج، بوزن خمسمئة نقطة لكل قسم. فمن الممكِّنات حملت القيادةُ ١٠٠، والسياسةُ والاستراتيجية ٨٠، والعاملون ٩٠، والشراكاتُ والموارد ٩٠، والعملياتُ ١٤٠. ومن النتائج حملت نتائجُ العملاء ٢٠٠، ونتائجُ الأداء الرئيسة ١٥٠، ونتائجُ العاملين ٩٠، ونتائجُ المجتمع ٦٠. وقيل إنّ الممكِّنات تقود النتائج وإنّ النتائج تُغذّي التعلّم راجعًا إلى الممكِّنات. ومن قرأ ورقةً تجريبية عن EFQM نُشرت قبل ٢٠٢٠ تقريبًا فإنما يقرأ عن هذه البِنية لا عن الحالية.
ورادار هو الجزء الذي يؤدّي العمل. والحروف تدلّ على النتائج والنهج والنشر والتقييم والتنقيح. ففي معيار الممكِّنات يقيّم المقيِّمُ هل النهجُ سليمٌ مندمجٌ مع الاستراتيجية، وهل نُشر نشرًا منهجيًّا في الأجزاء المعنيّة من المنظمة، وهل يُقاس ويُتعلَّم منه ويُحسَّن. وفي معيار النتائج يقيّم الملاءمةَ وقابلية الاستعمال، أي نطاق البيانات وتشريحَها وسلامتها، ثم الأداء، أي الاتجاهات عبر سنوات، وبلوغَ المستهدفات، والمقارنةَ بمرجعياتٍ خارجية ذات صلة، والثقةَ في استدامة المستوى.
ويُسجَّل كل عنصرٍ من رادار في نطاقاتٍ مئوية من صفرٍ إلى مئة، ثم تُجمع العناصر في نسبةٍ واحدة لكل معيار، وتُضرب تلك النسبة في وزن المعيار من النقاط. ثم تُعرَّف مستوياتُ التكريم بنطاقاتٍ من المجموع. والحسابُ بسيط، وهذا بعضُ جاذبيته، وهو يخفي أنّ كل مدخلٍ إليه حكمُ إنسانٍ مدرَّب على وثيقةٍ مكتوبة.
العلاقة ببالدريج وبالجوائز الوطنية. يشترك بالدريج وEFQM في البناء ويختلفان في أسماء المعايير وأوزانها ومفرداتها. فيضع بالدريج نحو ٤٥٠ من نقاطه الألف على النتائج في مقابل ٤٠٠ في بِنية EFQM لسنة ٢٠٢٠، وينظّم الممكِّنات في سبع فئاتٍ لا في خمسة معايير. وأكثرُ جوائز التميز الوطنية، في أوروبا وفيما وراءها بكثير، مشتقٌّ من أحدهما، مع إلحاق معاييرَ محلّية في العادة. وهذا الاشتقاق هو ما يجعل هذه الأسرة جديرةً بالدرس: الأداةُ نفسها، تطبّقها جهاتٌ مختلفة، على جماهيرَ مختلفة، بمعاييرَ ملحَقة مختلفة.
Decide first what kind of object EFQM is in the study. There are three honest choices. It is a measurement instrument whose structure and validity are being examined. It is an intervention whose effect on later performance is being estimated. Or it is the source of a dataset, because assessment generates scored, dated, comparable records of organizations that would otherwise be opaque. Each choice implies a different design. A study that does not choose usually ends up with the fourth option, which is citing the model as a framework and then not using it.
The instrument design. Take criterion-part scores from real assessments, not survey items written to resemble the criteria, and test the assumed structure with confirmatory factor analysis. Then test the weights: estimate the relation between each criterion and an outcome measured later, and compare the estimated importance with the committee weight. The published attempts to do this find substantial divergence, and the exercise has never been done on a non-European assessed population.
The intervention design. Score at time t, objective performance at t plus two or three years, on the full population of assessed organizations rather than the recognized ones, with controls for size, sector and prior performance. The award-winner event study is the other version, comparing recognized organizations with matched non-recognized ones on accounting performance around the recognition date, which is the design Hendricks and Singhal built for quality awards generally.
Designs that fail, and they are the common ones. Comparing award winners with the general population, which selects on the outcome. Correlating a self-assessment score with self-reported performance collected from the same respondent in the same questionnaire, which confounds the finding with common method variance and with the respondent's motive to justify the effort. Cross-sectional surveys of managers asked whether they agree that leadership drives results, which test nothing. And any longitudinal series that crosses a model revision without saying so.
Angles that the regional setting makes distinctive.
احسم أولًا أيّ شيءٍ يكون EFQM في الدراسة. فالخيارات الأمينة ثلاثة. إمّا أن يكون أداةَ قياسٍ تُفحَص بِنيتُها وصدقُها. وإمّا أن يكون تدخّلًا يُقدَّر أثرُه في الأداء اللاحق. وإمّا أن يكون مصدرًا لبيانات، إذ يولّد التقويمُ سجلّاتٍ مُدرَّجة مؤرَّخة قابلة للمقارنة عن منظماتٍ لولاه لكانت معتمة. ولكل خيارٍ تصميمٌ مختلف. والدراسةُ التي لا تختار تنتهي عادةً إلى الخيار الرابع، وهو الاستشهاد بالنموذج إطارًا ثم عدمُ استعماله.
تصميم الأداة. خذ درجات أجزاء المعايير من تقويماتٍ حقيقية لا من بنود مسحٍ كُتبت لتشبه المعايير، واختبر البِنية المفترَضة بالتحليل العاملي التوكيدي. ثم اختبر الأوزان: قدّر العلاقة بين كل معيارٍ ومخرَجٍ يُقاس لاحقًا، وقارن الأهمية المقدَّرة بوزن اللجنة. والمحاولاتُ المنشورة لهذا تجد تباعدًا كبيرًا، ولم يُجرَ هذا العملُ قطّ على جمهورٍ مقوَّم غير أوروبي.
تصميم التدخّل. الدرجةُ في الزمن الأول، والأداءُ الموضوعي بعد سنتين أو ثلاث، على جميع المنظمات المقوَّمة لا على المكرَّمة منها، مع ضبطٍ للحجم والقطاع والأداء السابق. ودراسةُ حدث الفوز هي الصورة الأخرى، بمقارنة المنظمات المكرَّمة بنظائرَ مطابَقة غير مكرَّمة في الأداء المحاسبي حول تاريخ التكريم، وهو التصميم الذي بناه هندريكس وسينغال لجوائز الجودة عمومًا.
والتصاميم التي تفشل، وهي الشائعة. مقارنةُ الفائزين بالجمهور العامّ، وهي انتقاءٌ على المخرَج. واقترانُ درجة التقويم الذاتي بأداءٍ يبلغه المستجيبُ نفسه في الاستبانة نفسها، وهذا يخلط النتيجة بتباين المنهج المشترك وبدافع المستجيب إلى تبرير الجهد. ومسوحٌ مقطعية تسأل المديرين هل يوافقون على أنّ القيادة تقود النتائج، وهي لا تختبر شيئًا. وكلُّ سلسلةٍ زمنية تعبر مراجعةً للنموذج من غير أن تذكر ذلك.
زوايا يجعلها السياق الإقليمي متميّزة.
It is an assessment framework, not a theory. It contains no proposition that could be shown false. It describes what a well-managed organization is said to look like and supplies a procedure for grading the resemblance. That is a useful thing to own and it is not a theoretical foundation. A thesis that names EFQM as its theory has no theory, and the gap usually shows in the hypotheses, which turn out to be restatements of the criteria rather than derivations from anything.
The weights are conventions set by committee, not estimates. Two hundred points for creating sustainable value and one hundred for engaging stakeholders express a negotiated judgement about relative importance. No estimation supports the ratio. The weights have changed at several revisions without any new evidence being cited for the change, and empirical attempts to recover criterion importance from data do not reproduce them. Any study that treats the total score as a meaningful composite has accepted the committee's judgement as a measurement decision.
Self-assessment and assessor training drive the scores. The same organization scores differently depending on who writes the application, how well the writer knows RADAR, and which assessor team reads it. Award bodies know this and manage it with training, calibration and team assessment. What they do not do is publish inter-assessor reliability, so the single most important psychometric property of the instrument is unavailable to anyone outside the process.
Studies linking scores to performance often select on the outcome. The convenient samples are the recognized organizations, and recognition is granted partly for performing well. A design comparing winners with non-applicants therefore cannot separate the model's effect from the selection that produced the sample. This is not a subtle flaw and it is present in a large share of the supportive literature, including work that is otherwise carefully executed.
Frequent revision defeats longitudinal comparison. A score from 1999, one from 2013 and one from 2020 are not the same quantity. The criterion set changed, the weights changed and the assessment logic was rewritten. Any trend line crossing a revision is measuring at least two things at once, and organizations that display improvement across a revision boundary are usually displaying an artifact.
The enabler and results ordering is read causally though it was never estimated as a causal model. The picture suggests that better enablers produce better results. Structural tests of that ordering return partial support with high collinearity among the enabler constructs, which means the model cannot distinguish the contribution of one enabler from another. That matters directly, because the practical advice a low score generates depends on which criterion is said to be weak.
هو إطارُ تقويمٍ لا نظرية. فليس فيه قضيةٌ يمكن إظهار كذبها. بل يصف ما يقال إنّ المنظمة حسنة الإدارة تبدو عليه، ويقدّم إجراءً لتدريج المشابهة. وذلك شيءٌ نافع يُملَك، وليس أساسًا نظريًّا. والرسالةُ التي تسمّي EFQM نظريتها لا نظرية لها، والفجوةُ تظهر عادةً في الفرضيات، إذ تتبيّن إعادةَ صياغةٍ للمعايير لا اشتقاقًا من شيء.
والأوزان أعرافٌ تضعها لجنة لا تقديرات. فمئتا نقطة لخلق قيمةٍ مستدامة ومئةٌ لإشراك أصحاب المصلحة تعبّران عن حكمٍ تفاوضيّ في الأهمية النسبية. ولا يسند هذه النسبةَ تقدير. وقد تغيّرت الأوزان في مراجعاتٍ عدّة من غير ذكر بيّنةٍ جديدة للتغيير، والمحاولاتُ التجريبية لاستخراج أهمية المعايير من البيانات لا تعيد إنتاجها. وكلُّ دراسةٍ تعامل المجموع مركَّبًا ذا معنًى فقد قبلت حكم اللجنة قرارًا في القياس.
والتقويم الذاتي وتدريب المقيِّمين هما ما يقود الدرجات. فالمنظمة نفسها تحصّل درجاتٍ مختلفة باختلاف كاتب الطلب، ومقدارِ معرفته برادار، وفريقِ المقيِّمين الذي يقرؤه. وجهاتُ الجوائز تعلم هذا وتديره بالتدريب والمعايرة والتقويم الجماعي. والذي لا تفعله نشرُ ثبات الاتّفاق بين المقيِّمين، فأهمُّ خاصّيةٍ قياسية في الأداة غيرُ متاحةٍ لأحدٍ خارج الإجراء.
والدراسات التي تربط الدرجات بالأداء كثيرًا ما تنتقي على المخرَج. فالعيّناتُ الميسّرة هي المنظمات المكرَّمة، والتكريمُ يُمنح جزئيًّا لحسن الأداء. فالتصميمُ الذي يقارن الفائزين بغير المتقدّمين عاجزٌ عن فصل أثر النموذج عن الانتقاء الذي أنتج العيّنة. وليس هذا عيبًا خفيًّا، وهو حاضرٌ في حصّةٍ كبيرة من الأدبيات المؤيِّدة، ومنها أعمالٌ مُتقنةٌ فيما عداه.
وكثرةُ المراجعة تُبطل المقارنة الطولية. فدرجةُ ١٩٩٩ ودرجةُ ٢٠١٣ ودرجةُ ٢٠٢٠ ليست كمّيةً واحدة. فقد تغيّرت مجموعةُ المعايير، وتغيّرت الأوزان، وأُعيدت كتابة منطق التقويم. وكلُّ خطّ اتجاهٍ يعبر مراجعةً يقيس شيئين على الأقلّ في آن، والمنظماتُ التي تعرض تحسّنًا عبر حدّ مراجعةٍ إنما تعرض أثرًا مصنوعًا في الغالب.
وترتيبُ الممكِّنات والنتائج يُقرأ سببيًّا وإن لم يُقدَّر نموذجًا سببيًّا قطّ. فالصورة توحي بأنّ تحسين الممكِّنات ينتج نتائجَ أحسن. والاختباراتُ البِنيوية لذلك الترتيب تعود بسندٍ جزئيّ مع تداخلٍ خطّي شديد بين بِناءات الممكِّنات، ومعناه أنّ النموذج لا يميّز إسهام ممكِّنٍ عن آخر. وهذا يمسّ الأمر مباشرةً، لأنّ النصيحة العملية التي تولّدها درجةٌ منخفضة تتوقّف على أيّ معيارٍ يقال إنّه ضعيف.