Saif Ali AlghamdiTransformation & Growth Advisor
تواصل
Business Fields and Theoriesحقول الأعمال ونظرياتهاQuality Management & Organizational Excellenceإدارة الجودة والتميز المؤسسي
QUALITY MANAGEMENT & ORGANIZATIONAL EXCELLENCE · PhDإدارة الجودة والتميز المؤسسي · دكتوراه

EFQM Excellence Modelنموذج التميز EFQM

SectionالقسمQuality Management & Organizational Excellenceإدارة الجودة والتميز المؤسسي
Reading timeزمن القراءة10 min١٠ دقيقة
ByإعدادSaif Alghamdiسيف الغامدي
One

Overview

Theory: EFQM Excellence Model
Primary field: Quality Management & Organizational Excellence
Core question: Can the way an organization is managed be assessed against a fixed criterion set and reduced to a score?
By: Saif Alghamdi

EFQM is an assessment framework, not a theory, and saying so at the start is not pedantry. A theory makes a claim that could be false. EFQM makes no such claim. It supplies a criterion set, a set of weights, a procedure for gathering evidence and a way of turning that evidence into a number, and its authority comes from adoption rather than from evidence.

What it actually contains is an architecture with four parts: a published list of criteria describing how a well-run organization behaves and what it achieves; fixed point weights that say how much each criterion counts; an assessment logic called RADAR that tells an assessor how to convert observed practice into a percentage; and a corps of trained assessors who apply it to a written application, usually with a site visit. The output is a score out of one thousand and a report naming strengths and areas for improvement.

Its usefulness is real and it is organizational rather than scientific. It gives a management team a shared vocabulary, forces the writing down of approaches that were tacit, and produces a comparison against a standard that is external to the organization. Those are genuine goods. None of them requires the model to be true, because there is nothing in it that could be true or false in the way a theory is.

For a research student the practical consequence is a warning. EFQM will not carry a thesis as its theoretical framework. It can serve as an instrument to be validated, as an intervention whose effect is estimated, or as a source of a dataset of assessed organizations. Those are three good uses. Presenting it as the theory the study rests on is the mistake the literature makes most often.

الأول

نظرة عامة

النظرية: نموذج التميز EFQM
الحقل الأساسي: إدارة الجودة والتميز المؤسسي
السؤال الجوهري: هل يمكن تقويم طريقة إدارة المنظمة بمعاييرَ ثابتة وردُّها إلى درجة؟
إعداد: سيف الغامدي

نموذج EFQM إطارُ تقويمٍ لا نظرية، وقولُ ذلك ابتداءً ليس تشدّقًا. فالنظرية تقول دعوى يمكن أن تكون كاذبة. وهذا النموذج لا يقول دعوى من هذا النوع. بل يقدّم مجموعةَ معايير، ومجموعةَ أوزان، وإجراءً لجمع الأدلّة، وطريقةً لتحويل الأدلّة إلى رقم، وسلطتُه آتيةٌ من الانتشار لا من البيّنة.

والذي يحويه فعلًا بناءٌ من أربعة أجزاء: قائمةٌ منشورة من المعايير تصف كيف تسلك المنظمة حسنةُ الإدارة وماذا تحقّق؛ وأوزانٌ ثابتة من النقاط تقول كم يزن كل معيار؛ ومنطقُ تقويمٍ يسمّى رادار يخبر المقيِّم كيف يحوّل الممارسة الملاحَظة إلى نسبة مئوية؛ وسلكٌ من المقيِّمين المدرَّبين يطبّقونه على طلبٍ مكتوب، مع زيارةٍ ميدانية في العادة. والمخرَجُ درجةٌ من ألف، وتقريرٌ يسمّي مواطن القوّة ومجالات التحسين.

ونفعُه حقيقيّ، وهو نفعٌ تنظيمي لا علميّ. فهو يعطي فريق الإدارة مفرداتٍ مشتركة، ويُلجئ إلى كتابة مناهجَ كانت ضمنية، وينتج مقارنةً بمعيارٍ خارجٍ عن المنظمة. وهذه خيراتٌ صادقة. ولا يقتضي شيءٌ منها أن يكون النموذج صادقًا، إذ ليس فيه ما يمكن أن يصدق أو يكذب كما تصدق النظرية أو تكذب.

واللازم العملي لطالب البحث تحذير. فـ EFQM لن يحمل رسالةً بوصفه إطارها النظري. وهو يصلح أداةً تُصدَّق، أو تدخّلًا يُقدَّر أثرُه، أو مصدرًا لبياناتٍ عن منظماتٍ مقوَّمة. وهذه ثلاثة استعمالاتٍ حسنة. وعرضُه نظريةً تقوم عليها الدراسة هو الخطأ الذي تقع فيه الأدبيات أكثر من غيره.

Two

Where It Came From

The Deming Prize, 1951. The oldest of the award-based quality assessments, established in Japan and named for a statistician whose lectures on process control were taken more seriously there than at home. It set the pattern that everything after it followed: an application document, an external examination, and public recognition tied to management practice rather than to a product test.

Baldrige, 1987. The Malcolm Baldrige National Quality Award in the United States built the architecture EFQM inherited almost intact. A published criterion set with fixed point weights, a written application, trained volunteer examiners, a site visit, a score out of one thousand, and a national award presented at the top of government. The design choice that mattered most was making the criteria public, which turned the award into a self-assessment instrument used by thousands of organizations that never applied.

The founding, 1988 to 1991. Fourteen European companies established the European Foundation for Quality Management in 1988. The model and the European Quality Award followed in 1991, first awarded in 1992. The motivation was competitive rather than scholarly: European manufacturers wanted an answer to Japanese quality practice and a European counterpart to the American award.

The 1999 revision and the revisions after it. The 1999 version renamed the framework the EFQM Excellence Model, introduced RADAR as the assessment logic in place of the earlier scoring matrix, and settled the nine criteria into five enablers and four results weighted five hundred points against five hundred. Further revisions followed in 2003, 2010 and 2013, adjusting criterion parts, the fundamental concepts and the weights. None of the reweightings was published alongside an estimation that would justify it.

The 2020 restructure. The nine-box picture was abandoned. The model was rebuilt around three blocks, Direction, Execution and Results, holding seven criteria, with a stronger emphasis on purpose, transformation and ecosystem thinking. RADAR was retained and rewritten. The change was substantial enough that scores under the new structure are not comparable with scores under the old one, which is rarely stated when trend charts are drawn.

Bou-Llusar and colleagues, 2009. The most serious empirical examination, testing the enabler and results structure with structural equation models on assessment data and comparing the model against the Baldrige criteria. The support was partial. The enabler constructs were highly collinear, the assumed separation between criteria did not reproduce cleanly, and the fit of the postulated causal ordering was weaker than its universal use in practice implies.

الثاني

الأصل والنشأة

جائزة ديمنغ، ١٩٥١. أقدمُ تقويمات الجودة القائمة على الجوائز، أُنشئت في اليابان وسُمّيت باسم إحصائيٍّ أُخذت محاضراتُه في ضبط العمليات هناك مأخذًا أجدَّ مما أُخذت في بلده. وقد أرست النمط الذي تبعه كل ما جاء بعدها: وثيقةُ طلب، وفحصٌ خارجي، وتكريمٌ علنيّ معلَّق بممارسة الإدارة لا باختبار منتج.

بالدريج، ١٩٨٧. بنت جائزةُ مالكوم بالدريج الوطنية للجودة في الولايات المتحدة البناءَ الذي ورثه EFQM كاملًا تقريبًا. مجموعةُ معايير منشورة بأوزانٍ ثابتة من النقاط، وطلبٌ مكتوب، وفاحصون متطوّعون مدرَّبون، وزيارةٌ ميدانية، ودرجةٌ من ألف، وجائزةٌ وطنية تُسلَّم في أعلى الدولة. وأهمّ اختيارٍ تصميمي فيها نشرُ المعايير، وهو ما حوّل الجائزة إلى أداة تقويمٍ ذاتي تستعملها آلافُ المنظمات التي لم تتقدّم إليها قطّ.

التأسيس، ١٩٨٨ إلى ١٩٩١. أنشأت أربع عشرة شركة أوروبية المؤسسةَ الأوروبية لإدارة الجودة سنة ١٩٨٨. وتبعها النموذجُ والجائزةُ الأوروبية للجودة سنة ١٩٩١، ومُنحت أول مرّة سنة ١٩٩٢. وكان الباعث تنافسيًّا لا علميًّا: أراد الصنّاع الأوروبيون جوابًا عن الممارسة اليابانية في الجودة، ونظيرًا أوروبيًّا للجائزة الأمريكية.

مراجعة ١٩٩٩ وما بعدها. سمّت نسخةُ ١٩٩٩ الإطارَ نموذج التميز EFQM، وأدخلت رادار منطقًا للتقويم بدل مصفوفة التسجيل السابقة، واستقرّت بالمعايير التسعة على خمسة ممكِّنات وأربع نتائج بوزن خمسمئة نقطة في مقابل خمسمئة. وتلتها مراجعاتٌ سنة ٢٠٠٣ و٢٠١٠ و٢٠١٣، عدّلت أجزاء المعايير والمفاهيم الأساسية والأوزان. ولم تُنشر واحدةٌ من إعادات الترجيح مقرونةً بتقديرٍ يسوّغها.

إعادة بناء ٢٠٢٠. هُجرت صورةُ الصناديق التسعة. وأُعيد بناء النموذج حول ثلاث كتل: التوجّه والتنفيذ والنتائج، تحوي سبعة معايير، مع تشديدٍ أقوى على الغاية والتحوّل والتفكير المنظومي. واحتُفظ برادار بعد إعادة صياغته. وكان التغيير من الجوهرية بحيث لا تقارَن درجاتُ البِنية الجديدة بدرجات القديمة، وقلّما يُذكر ذلك حين تُرسم منحنياتُ الاتجاه.

بو لوسار وزملاؤه، ٢٠٠٩. أجدُّ فحصٍ تجريبي، اختبر بِنية الممكِّنات والنتائج بنماذج المعادلات البِنيوية على بيانات تقويمٍ فعلية، وقارن النموذج بمعايير بالدريج. وجاء السند جزئيًّا. فبِناءات الممكِّنات كانت شديدة التداخل الخطّي، والفصلُ المفترَض بين المعايير لم يظهر نظيفًا، وملاءمةُ الترتيب السببي المفترَض كانت أضعف مما يوحي به استعمالُه الشامل في الممارسة.

Three

How It Works

The architecture. Four parts working together: criteria that describe practice and achievement, weights that price each criterion, an assessment logic that converts observation into a percentage, and assessors trained to apply that logic consistently. Remove any one and the score stops meaning anything. Most published discussion of EFQM addresses only the first part, which is why so much of it misses where the score actually comes from.

BlockCriterion in the 2020 structurePoints
DirectionPurpose, vision and strategy100
DirectionOrganisational culture and leadership100
ExecutionEngaging stakeholders100
ExecutionCreating sustainable value200
ExecutionDriving performance and transformation100
ResultsStakeholder perceptions200
ResultsStrategic and operational performance200

The older arrangement is still the one most literature was written about. From 1999 the model held nine criteria split into five enablers and four results, weighted five hundred points each. Among the enablers, leadership carried 100, policy and strategy 80, people 90, partnerships and resources 90, and processes 140. Among the results, customer results carried 200, key performance results 150, people results 90 and society results 60. The enablers were said to drive the results and the results were said to feed learning back into the enablers. Anyone reading an empirical paper on EFQM published before roughly 2020 is reading about this structure, not the current one.

RADAR is the part that does the work. The letters stand for Results, Approach, Deploy, Assess and Refine. For an enabler criterion the assessor scores whether the approach is sound and integrated with the strategy, whether it is deployed systematically across the relevant parts of the organization, and whether it is measured, learned from and improved. For a results criterion the assessor scores relevance and usability, meaning scope, segmentation and integrity of the data, and performance, meaning trends over several years, achievement against targets, comparison with relevant external benchmarks, and confidence that the level will be sustained.

criterion score = RADAR percentage × criterion points
total = sum of the seven criterion scores, out of 1,000

Each RADAR element is scored in percentage bands from zero to one hundred, the elements are combined into one percentage per criterion, and that percentage is multiplied by the criterion's point weight. Recognition tiers are then defined by bands of the total. The arithmetic is simple, which is part of the appeal, and it hides the fact that every input to it is a trained human judgement about a written document.

The score measures evidenced management approach, not quality. RADAR rewards an approach that is written down, deployed visibly and reviewed on a cycle. An organization that runs well by habit and documents nothing scores badly; an organization that documents thoroughly and performs adequately scores well. This is a defensible thing to measure, and it is not what most readers of a score believe they are reading. The related error is treating the enabler and results split as a causal model. EFQM presents a direction of influence, but the weights and the arrangement were set by committee, and the causal claim has been tested rather than assumed only in a handful of studies.

Relation to Baldrige and to national awards. Baldrige and EFQM share their architecture and differ in their criterion names, their weights and their vocabulary. Baldrige places roughly 450 of its 1,000 points on results against 400 in the 2020 EFQM structure, and organizes enablers into seven categories rather than five criteria. Most national excellence awards, in Europe and well beyond it, are derived from one or the other, usually with local criteria appended. That derivation is what makes the family worth studying: the same instrument, applied by different bodies, to different populations, with different appended criteria.

الثالث

الآلية والبنية

البناء. أربعةُ أجزاء تعمل معًا: معايير تصف الممارسة والإنجاز، وأوزانٌ تسعّر كل معيار، ومنطقُ تقويمٍ يحوّل الملاحظة إلى نسبة، ومقيِّمون مدرَّبون على تطبيق ذلك المنطق باتّساق. وارفع واحدًا منها تتوقّف الدرجة عن أن تعني شيئًا. وأكثرُ ما يُنشر في EFQM يعالج الجزء الأول وحده، ولهذا يفوت كثيرًا منه من أين تأتي الدرجة فعلًا.

الكتلةالمعيار في بِنية ٢٠٢٠النقاط
التوجّهالغاية والرؤية والاستراتيجية١٠٠
التوجّهالثقافة المؤسسية والقيادة١٠٠
التنفيذإشراك أصحاب المصلحة١٠٠
التنفيذخلق قيمة مستدامة٢٠٠
التنفيذدفع الأداء والتحوّل١٠٠
النتائجتصوّرات أصحاب المصلحة٢٠٠
النتائجالأداء الاستراتيجي والتشغيلي٢٠٠

والترتيب الأقدم هو الذي كُتبت عنه أكثرُ الأدبيات. فمنذ ١٩٩٩ حمل النموذج تسعة معايير مقسومةً إلى خمسة ممكِّنات وأربع نتائج، بوزن خمسمئة نقطة لكل قسم. فمن الممكِّنات حملت القيادةُ ١٠٠، والسياسةُ والاستراتيجية ٨٠، والعاملون ٩٠، والشراكاتُ والموارد ٩٠، والعملياتُ ١٤٠. ومن النتائج حملت نتائجُ العملاء ٢٠٠، ونتائجُ الأداء الرئيسة ١٥٠، ونتائجُ العاملين ٩٠، ونتائجُ المجتمع ٦٠. وقيل إنّ الممكِّنات تقود النتائج وإنّ النتائج تُغذّي التعلّم راجعًا إلى الممكِّنات. ومن قرأ ورقةً تجريبية عن EFQM نُشرت قبل ٢٠٢٠ تقريبًا فإنما يقرأ عن هذه البِنية لا عن الحالية.

ورادار هو الجزء الذي يؤدّي العمل. والحروف تدلّ على النتائج والنهج والنشر والتقييم والتنقيح. ففي معيار الممكِّنات يقيّم المقيِّمُ هل النهجُ سليمٌ مندمجٌ مع الاستراتيجية، وهل نُشر نشرًا منهجيًّا في الأجزاء المعنيّة من المنظمة، وهل يُقاس ويُتعلَّم منه ويُحسَّن. وفي معيار النتائج يقيّم الملاءمةَ وقابلية الاستعمال، أي نطاق البيانات وتشريحَها وسلامتها، ثم الأداء، أي الاتجاهات عبر سنوات، وبلوغَ المستهدفات، والمقارنةَ بمرجعياتٍ خارجية ذات صلة، والثقةَ في استدامة المستوى.

درجة المعيار = نسبة رادار المئوية × نقاط المعيار
المجموع = حاصل جمع درجات المعايير السبعة، من ١٠٠٠

ويُسجَّل كل عنصرٍ من رادار في نطاقاتٍ مئوية من صفرٍ إلى مئة، ثم تُجمع العناصر في نسبةٍ واحدة لكل معيار، وتُضرب تلك النسبة في وزن المعيار من النقاط. ثم تُعرَّف مستوياتُ التكريم بنطاقاتٍ من المجموع. والحسابُ بسيط، وهذا بعضُ جاذبيته، وهو يخفي أنّ كل مدخلٍ إليه حكمُ إنسانٍ مدرَّب على وثيقةٍ مكتوبة.

الدرجة تقيس نهجَ إدارةٍ مُدلَّلًا عليه لا تقيس الجودة. فرادار يكافئ نهجًا مكتوبًا منشورًا ظاهرًا مراجَعًا على دورة. فالمنظمةُ التي تُحسِن بالعادة ولا توثّق شيئًا تحصّل درجةً رديئة؛ والمنظمةُ التي توثّق توثيقًا محكمًا وتؤدّي أداءً كافيًا تحصّل درجةً حسنة. وهذا شيءٌ يُدافَع عن قياسه، وليس هو ما يظنّ أكثرُ قرّاء الدرجة أنهم يقرؤونه. والخطأ المصاحب معاملةُ قسمة الممكِّنات والنتائج نموذجًا سببيًّا. فالنموذج يعرض اتجاه تأثير، لكنّ الأوزان والترتيب وُضعا بلجنة، ولم تُختبر الدعوى السببية بدل افتراضها إلا في دراساتٍ معدودة.

العلاقة ببالدريج وبالجوائز الوطنية. يشترك بالدريج وEFQM في البناء ويختلفان في أسماء المعايير وأوزانها ومفرداتها. فيضع بالدريج نحو ٤٥٠ من نقاطه الألف على النتائج في مقابل ٤٠٠ في بِنية EFQM لسنة ٢٠٢٠، وينظّم الممكِّنات في سبع فئاتٍ لا في خمسة معايير. وأكثرُ جوائز التميز الوطنية، في أوروبا وفيما وراءها بكثير، مشتقٌّ من أحدهما، مع إلحاق معاييرَ محلّية في العادة. وهذا الاشتقاق هو ما يجعل هذه الأسرة جديرةً بالدرس: الأداةُ نفسها، تطبّقها جهاتٌ مختلفة، على جماهيرَ مختلفة، بمعاييرَ ملحَقة مختلفة.

Four

Using It in Research

Decide first what kind of object EFQM is in the study. There are three honest choices. It is a measurement instrument whose structure and validity are being examined. It is an intervention whose effect on later performance is being estimated. Or it is the source of a dataset, because assessment generates scored, dated, comparable records of organizations that would otherwise be opaque. Each choice implies a different design. A study that does not choose usually ends up with the fourth option, which is citing the model as a framework and then not using it.

The instrument design. Take criterion-part scores from real assessments, not survey items written to resemble the criteria, and test the assumed structure with confirmatory factor analysis. Then test the weights: estimate the relation between each criterion and an outcome measured later, and compare the estimated importance with the committee weight. The published attempts to do this find substantial divergence, and the exercise has never been done on a non-European assessed population.

The intervention design. Score at time t, objective performance at t plus two or three years, on the full population of assessed organizations rather than the recognized ones, with controls for size, sector and prior performance. The award-winner event study is the other version, comparing recognized organizations with matched non-recognized ones on accounting performance around the recognition date, which is the design Hendricks and Singhal built for quality awards generally.

Designs that fail, and they are the common ones. Comparing award winners with the general population, which selects on the outcome. Correlating a self-assessment score with self-reported performance collected from the same respondent in the same questionnaire, which confounds the finding with common method variance and with the respondent's motive to justify the effort. Cross-sectional surveys of managers asked whether they agree that leadership drives results, which test nothing. And any longitudinal series that crosses a model revision without saying so.

Almost every positive finding about EFQM selects on the outcome. Award winners were chosen partly because they perform well, so their subsequent performance is not evidence that the model caused anything. The design that survives review draws its sample frame from all assessed organizations, keeps the score as a continuous variable including the low end, and measures performance strictly after the assessment date. If the low scorers are not in the sample, the study cannot answer the question it asks.

Angles that the regional setting makes distinctive.

  • National excellence awards as a public dataset. Where a national award operates on EFQM-derived criteria and publishes the roster of participating and recognized entities across several cycles, the result is an unusual thing: a dated, externally assessed, comparable record covering organizations that publish nothing else. That includes public bodies with no share price, no analyst coverage and no annual report worth the name. For those organizations the assessment record may be the only comparable performance data that exists.
  • The state as both assessor and assessed. Where government excellence programmes assess ministries, agencies and municipalities, and the assessing body is itself part of government, the assessment is a governance instrument as much as a measurement one. The scores then carry budget and career consequences, which is precisely the condition under which a measure degrades. Testing for score inflation over successive cycles is direct and feasible.
  • Appended local criteria as a natural experiment on weights. Where a national programme adds criteria for workforce localization or alignment with a national transformation programme to the imported EFQM base, the appended and imported criteria can be compared on how well each predicts later objective outcomes. That is a clean test of whether committee weights travel across institutional settings.
  • Assessor variance as an estimable quantity. Where a national body trains, certifies and records its assessor teams, the share of score variance attributable to the assessing team rather than the applicant can be estimated with a cross-classified model. The question is fundamental to every use of the score, it is answerable from records the award bodies already hold, and it appears to be unpublished anywhere.
  • Bilingual assessment. Applications written in Arabic are assessed against a criterion set drafted in English, and the RADAR vocabulary of approach, deployment, assessment and refinement has no settled Arabic rendering. Whether the same evidence receives the same score in the two languages is measurable, and it bears on every regional score ever awarded.
  • Young organizations and the trend requirement. RADAR asks for results trends over several years and confidence in future performance. Entities created recently inside large transformation programmes have no such history, and rapidly expanding capacity breaks the trend logic even where history exists. How assessors handle that gap, and whether it systematically depresses or inflates scores for new entities, is an open question with a ready sample.
الرابع

التوظيف البحثي

احسم أولًا أيّ شيءٍ يكون EFQM في الدراسة. فالخيارات الأمينة ثلاثة. إمّا أن يكون أداةَ قياسٍ تُفحَص بِنيتُها وصدقُها. وإمّا أن يكون تدخّلًا يُقدَّر أثرُه في الأداء اللاحق. وإمّا أن يكون مصدرًا لبيانات، إذ يولّد التقويمُ سجلّاتٍ مُدرَّجة مؤرَّخة قابلة للمقارنة عن منظماتٍ لولاه لكانت معتمة. ولكل خيارٍ تصميمٌ مختلف. والدراسةُ التي لا تختار تنتهي عادةً إلى الخيار الرابع، وهو الاستشهاد بالنموذج إطارًا ثم عدمُ استعماله.

تصميم الأداة. خذ درجات أجزاء المعايير من تقويماتٍ حقيقية لا من بنود مسحٍ كُتبت لتشبه المعايير، واختبر البِنية المفترَضة بالتحليل العاملي التوكيدي. ثم اختبر الأوزان: قدّر العلاقة بين كل معيارٍ ومخرَجٍ يُقاس لاحقًا، وقارن الأهمية المقدَّرة بوزن اللجنة. والمحاولاتُ المنشورة لهذا تجد تباعدًا كبيرًا، ولم يُجرَ هذا العملُ قطّ على جمهورٍ مقوَّم غير أوروبي.

تصميم التدخّل. الدرجةُ في الزمن الأول، والأداءُ الموضوعي بعد سنتين أو ثلاث، على جميع المنظمات المقوَّمة لا على المكرَّمة منها، مع ضبطٍ للحجم والقطاع والأداء السابق. ودراسةُ حدث الفوز هي الصورة الأخرى، بمقارنة المنظمات المكرَّمة بنظائرَ مطابَقة غير مكرَّمة في الأداء المحاسبي حول تاريخ التكريم، وهو التصميم الذي بناه هندريكس وسينغال لجوائز الجودة عمومًا.

والتصاميم التي تفشل، وهي الشائعة. مقارنةُ الفائزين بالجمهور العامّ، وهي انتقاءٌ على المخرَج. واقترانُ درجة التقويم الذاتي بأداءٍ يبلغه المستجيبُ نفسه في الاستبانة نفسها، وهذا يخلط النتيجة بتباين المنهج المشترك وبدافع المستجيب إلى تبرير الجهد. ومسوحٌ مقطعية تسأل المديرين هل يوافقون على أنّ القيادة تقود النتائج، وهي لا تختبر شيئًا. وكلُّ سلسلةٍ زمنية تعبر مراجعةً للنموذج من غير أن تذكر ذلك.

تكاد كل نتيجةٍ إيجابية عن EFQM تنتقي على المخرَج. فالفائزون اختيروا جزئيًّا لحسن أدائهم، فأداؤهم اللاحق ليس دليلًا على أنّ النموذج سبّب شيئًا. والتصميم الذي ينجو من التحكيم يسحب إطار عيّنته من جميع المنظمات المقوَّمة، ويُبقي الدرجة متغيّرًا متّصلًا يشمل طرفها الأدنى، ويقيس الأداء بعد تاريخ التقويم قطعًا. وإن لم يكن أصحابُ الدرجات المنخفضة في العيّنة فالدراسة عاجزةٌ عن جواب سؤالها.

زوايا يجعلها السياق الإقليمي متميّزة.

  • جوائز التميز الوطنية بياناتٍ عامّة. فحيث تعمل جائزةٌ وطنية بمعاييرَ مشتقّة من EFQM وتنشر قائمة الجهات المشاركة والمكرَّمة عبر دوراتٍ عدّة، نتج شيءٌ غير معتاد: سجلٌّ مؤرَّخ مقوَّمٌ من خارجٍ قابلٌ للمقارنة يغطّي منظماتٍ لا تنشر سواه. ومن ذلك جهاتٌ عامّة بلا سعر سهمٍ ولا تغطيةِ محلّلين ولا تقريرٍ سنويّ يستحقّ الاسم. ولهذه المنظمات قد يكون سجلُّ التقويم بيانَ الأداء المقارَن الوحيد الموجود.
  • الدولة مقيِّمًا ومقوَّمًا معًا. فحيث تقوّم برامجُ التميز الحكومية الوزاراتِ والهيئاتِ والبلديات، وتكون الجهةُ المقيِّمة جزءًا من الحكومة نفسها، صار التقويم أداةَ حوكمةٍ بقدر ما هو أداةُ قياس. وتحمل الدرجاتُ حينئذٍ لوازمَ في الميزانية والمسار الوظيفي، وذلك بعينه الشرط الذي يفسد المقياسُ تحته. واختبارُ تضخّم الدرجات عبر الدورات المتعاقبة مباشرٌ وممكن.
  • المعايير المحلّية الملحَقة تجربةً طبيعية في الأوزان. فحيث يضيف برنامجٌ وطني معاييرَ لتوطين الوظائف أو للمواءمة مع برنامج تحوّلٍ وطني إلى الأساس المستورَد، أمكنت مقارنةُ المعايير الملحَقة بالمستورَدة في مقدار تنبّؤ كلٍّ منها بمخرجاتٍ موضوعية لاحقة. وذلك اختبارٌ نظيف لهل تنتقل أوزانُ اللجان بين السياقات المؤسسية.
  • تباينُ المقيِّمين كمّيةً قابلة للتقدير. فحيث تدرّب جهةٌ وطنية فرقَ مقيِّميها وتعتمدهم وتسجّلهم، أمكن تقديرُ حصّة تباين الدرجة العائدة إلى فريق التقويم لا إلى مقدّم الطلب بنموذجٍ متعدّد التصنيف. والسؤال أساسيٌّ لكل استعمالٍ للدرجة، وجوابُه متاحٌ من سجلّاتٍ تحتفظ بها جهاتُ الجوائز أصلًا، ولا يبدو أنه نُشر في مكان.
  • التقويم ثنائي اللغة. فالطلبات المكتوبة بالعربية تُقوَّم بمجموعة معاييرَ صيغت بالإنجليزية، ولمفردات رادار من نهجٍ ونشرٍ وتقييمٍ وتنقيح ليس في العربية مقابلٌ مستقرّ. وهل ينال الدليلُ نفسه الدرجةَ نفسها باللغتين سؤالٌ قابل للقياس، وهو يمسّ كل درجةٍ مُنحت في الإقليم.
  • المنظمات الفتيّة ومتطلَّب الاتجاه. فرادار يطلب اتجاهاتِ نتائجَ عبر سنوات وثقةً في الأداء المستقبلي. والكياناتُ المنشأة حديثًا داخل برامج تحوّلٍ كبرى لا تاريخ لها من هذا، والتوسّعُ السريع في الطاقة يكسر منطق الاتجاه ولو وُجد التاريخ. وكيف يعالج المقيِّمون تلك الفجوة، وهل تخفض درجات الكيانات الجديدة منهجيًّا أو ترفعها، سؤالٌ مفتوح بعيّنةٍ جاهزة.
Five

Limits and Critique

It is an assessment framework, not a theory. It contains no proposition that could be shown false. It describes what a well-managed organization is said to look like and supplies a procedure for grading the resemblance. That is a useful thing to own and it is not a theoretical foundation. A thesis that names EFQM as its theory has no theory, and the gap usually shows in the hypotheses, which turn out to be restatements of the criteria rather than derivations from anything.

The weights are conventions set by committee, not estimates. Two hundred points for creating sustainable value and one hundred for engaging stakeholders express a negotiated judgement about relative importance. No estimation supports the ratio. The weights have changed at several revisions without any new evidence being cited for the change, and empirical attempts to recover criterion importance from data do not reproduce them. Any study that treats the total score as a meaningful composite has accepted the committee's judgement as a measurement decision.

Self-assessment and assessor training drive the scores. The same organization scores differently depending on who writes the application, how well the writer knows RADAR, and which assessor team reads it. Award bodies know this and manage it with training, calibration and team assessment. What they do not do is publish inter-assessor reliability, so the single most important psychometric property of the instrument is unavailable to anyone outside the process.

Studies linking scores to performance often select on the outcome. The convenient samples are the recognized organizations, and recognition is granted partly for performing well. A design comparing winners with non-applicants therefore cannot separate the model's effect from the selection that produced the sample. This is not a subtle flaw and it is present in a large share of the supportive literature, including work that is otherwise carefully executed.

Frequent revision defeats longitudinal comparison. A score from 1999, one from 2013 and one from 2020 are not the same quantity. The criterion set changed, the weights changed and the assessment logic was rewritten. Any trend line crossing a revision is measuring at least two things at once, and organizations that display improvement across a revision boundary are usually displaying an artifact.

The enabler and results ordering is read causally though it was never estimated as a causal model. The picture suggests that better enablers produce better results. Structural tests of that ordering return partial support with high collinearity among the enabler constructs, which means the model cannot distinguish the contribution of one enabler from another. That matters directly, because the practical advice a low score generates depends on which criterion is said to be weak.

Treat EFQM as an instrument or an intervention and never as the theory a study rests on. If the instrument is the subject, work from real assessment records rather than survey items written to resemble the criteria, and test the weights instead of accepting them. If the effect is the subject, keep the low scorers in the sample and measure outcomes strictly after the assessment date. Do not compare scores across a model revision. And where a national award assesses organizations that publish nothing else, recognize what is actually on offer: a dated external assessment of bodies that are otherwise invisible to research, which is worth more than one more study of the winners.
الخامس

الحدود والنقد

هو إطارُ تقويمٍ لا نظرية. فليس فيه قضيةٌ يمكن إظهار كذبها. بل يصف ما يقال إنّ المنظمة حسنة الإدارة تبدو عليه، ويقدّم إجراءً لتدريج المشابهة. وذلك شيءٌ نافع يُملَك، وليس أساسًا نظريًّا. والرسالةُ التي تسمّي EFQM نظريتها لا نظرية لها، والفجوةُ تظهر عادةً في الفرضيات، إذ تتبيّن إعادةَ صياغةٍ للمعايير لا اشتقاقًا من شيء.

والأوزان أعرافٌ تضعها لجنة لا تقديرات. فمئتا نقطة لخلق قيمةٍ مستدامة ومئةٌ لإشراك أصحاب المصلحة تعبّران عن حكمٍ تفاوضيّ في الأهمية النسبية. ولا يسند هذه النسبةَ تقدير. وقد تغيّرت الأوزان في مراجعاتٍ عدّة من غير ذكر بيّنةٍ جديدة للتغيير، والمحاولاتُ التجريبية لاستخراج أهمية المعايير من البيانات لا تعيد إنتاجها. وكلُّ دراسةٍ تعامل المجموع مركَّبًا ذا معنًى فقد قبلت حكم اللجنة قرارًا في القياس.

والتقويم الذاتي وتدريب المقيِّمين هما ما يقود الدرجات. فالمنظمة نفسها تحصّل درجاتٍ مختلفة باختلاف كاتب الطلب، ومقدارِ معرفته برادار، وفريقِ المقيِّمين الذي يقرؤه. وجهاتُ الجوائز تعلم هذا وتديره بالتدريب والمعايرة والتقويم الجماعي. والذي لا تفعله نشرُ ثبات الاتّفاق بين المقيِّمين، فأهمُّ خاصّيةٍ قياسية في الأداة غيرُ متاحةٍ لأحدٍ خارج الإجراء.

والدراسات التي تربط الدرجات بالأداء كثيرًا ما تنتقي على المخرَج. فالعيّناتُ الميسّرة هي المنظمات المكرَّمة، والتكريمُ يُمنح جزئيًّا لحسن الأداء. فالتصميمُ الذي يقارن الفائزين بغير المتقدّمين عاجزٌ عن فصل أثر النموذج عن الانتقاء الذي أنتج العيّنة. وليس هذا عيبًا خفيًّا، وهو حاضرٌ في حصّةٍ كبيرة من الأدبيات المؤيِّدة، ومنها أعمالٌ مُتقنةٌ فيما عداه.

وكثرةُ المراجعة تُبطل المقارنة الطولية. فدرجةُ ١٩٩٩ ودرجةُ ٢٠١٣ ودرجةُ ٢٠٢٠ ليست كمّيةً واحدة. فقد تغيّرت مجموعةُ المعايير، وتغيّرت الأوزان، وأُعيدت كتابة منطق التقويم. وكلُّ خطّ اتجاهٍ يعبر مراجعةً يقيس شيئين على الأقلّ في آن، والمنظماتُ التي تعرض تحسّنًا عبر حدّ مراجعةٍ إنما تعرض أثرًا مصنوعًا في الغالب.

وترتيبُ الممكِّنات والنتائج يُقرأ سببيًّا وإن لم يُقدَّر نموذجًا سببيًّا قطّ. فالصورة توحي بأنّ تحسين الممكِّنات ينتج نتائجَ أحسن. والاختباراتُ البِنيوية لذلك الترتيب تعود بسندٍ جزئيّ مع تداخلٍ خطّي شديد بين بِناءات الممكِّنات، ومعناه أنّ النموذج لا يميّز إسهام ممكِّنٍ عن آخر. وهذا يمسّ الأمر مباشرةً، لأنّ النصيحة العملية التي تولّدها درجةٌ منخفضة تتوقّف على أيّ معيارٍ يقال إنّه ضعيف.

عامِل EFQM أداةً أو تدخّلًا ولا تعامله قطّ نظريةً تقوم عليها دراسة. فإن كانت الأداةُ هي الموضوع فاعمل من سجلّات تقويمٍ حقيقية لا من بنود مسحٍ كُتبت لتشبه المعايير، واختبر الأوزان بدل قبولها. وإن كان الأثرُ هو الموضوع فأبقِ أصحاب الدرجات المنخفضة في العيّنة، وقِس المخرجات بعد تاريخ التقويم قطعًا. ولا تقارن درجاتٍ عبر مراجعةٍ للنموذج. وحيث تقوّم جائزةٌ وطنية منظماتٍ لا تنشر سواها فاعرف ما المعروض فعلًا: تقويمٌ خارجيّ مؤرَّخ لجهاتٍ هي فيما عداه غيرُ مرئيةٍ للبحث، وذلك أثمن من دراسةٍ أخرى عن الفائزين.