SERVQUAL exists because service quality cannot be measured the way product quality is. A product can be tested against a specification before it reaches the customer. A service is produced and consumed at the same moment, varies with who delivers it, and has few physical attributes to inspect, so quality has to be defined by the customer's judgment rather than by conformance.
The instrument's definition is a difference: quality is perceived performance minus expectation. That formulation has consequences that are easy to overlook. A service can be objectively excellent and score badly because expectations were higher still, and a modest service can score well by exceeding low expectations. Managing quality therefore includes managing what the customer was led to expect.
SERVQUAL is the most widely used instrument in services research and also one of the most criticized, and a doctoral text should present both, since much applied work cites the instrument without acknowledging that its difference-score design has been seriously challenged.
وُجد SERVQUAL لأن جودة الخدمة لا تُقاس كما تُقاس جودة المنتج. فالمنتج يمكن اختباره على مواصفةٍ قبل بلوغه العميل. والخدمة تُنتَج وتُستهلك في اللحظة نفسها، وتتباين بمن يقدّمها، وقليلةُ الخصائص المادية القابلة للفحص، فوجب تعريف الجودة بحكم العميل لا بالمطابقة.
وتعريف الأداة فرق: الجودة أداءٌ مدرَك ناقصَ توقّع. ولهذه الصياغة نتائج يسهل إغفالها. فقد تكون خدمةٌ ممتازة موضوعيًّا فتنال درجةً سيئة لأن التوقّعات كانت أعلى، وقد تنال خدمةٌ متواضعة درجةً حسنة بتجاوزها توقّعاتٍ منخفضة. فإدارة الجودة إذن تشمل إدارة ما حُمل العميل على توقّعه.
وSERVQUAL أوسع الأدوات استعمالًا في بحوث الخدمات وأكثرها نقدًا كذلك، وينبغي لنصّ الدكتوراه أن يعرض الأمرين، فكثيرٌ من العمل التطبيقي يستشهد بالأداة من غير الإقرار بأن تصميمها القائم على درجة الفرق قد وُوجه بتحدٍّ جادّ.
Parasuraman, Zeithaml and Berry, 1985. The conceptual paper that came first. Exploratory work with service executives and customer focus groups produced the gaps model of service quality and an initial list of ten determinants of quality. The instrument came out of this, not the other way round, which is why it is worth reading the conceptual paper before the scale.
Parasuraman, Zeithaml and Berry, 1988. The instrument. The ten determinants were reduced through factor analysis to five dimensions and twenty two paired items, one asking what an excellent firm in the industry should offer and one asking what this firm delivered. The score for each item is the difference, and the dimension score is the average.
Cronin and Taylor, 1992. The most important challenge, and it is empirical rather than conceptual. They argue that service quality is better measured as performance alone, an instrument they call SERVPERF, and they show that the performance-only version explains more variance in overall quality and in purchase intention than the difference score does. Their statistical case against difference scores is strong: subtracting two measured variables carries the error of both and typically has lower reliability than either.
Parasuraman, Zeithaml and Berry, 1994. The reply, which concedes ground on the statistics and defends the expectations component on diagnostic grounds: performance alone tells a manager the score, while the gap tells them where the shortfall is. The paper also introduces a zone of tolerance between desired and adequate service, which is a genuine improvement on a single expectation standard.
Buttle, 1996. The consolidated critique, and the standard citation when acknowledging the instrument's problems.
باراسورامان وزيثامل وبيري، ١٩٨٥. الورقة المفاهيمية التي جاءت أولًا. أنتج عملٌ استكشافي مع مديري خدماتٍ ومجموعات تركيزٍ من العملاء نموذجَ الفجوات في جودة الخدمة وقائمةً أولية بعشرة محدّدات للجودة. وخرجت الأداة من هذا لا العكس، ولهذا حسُن قراءة الورقة المفاهيمية قبل المقياس.
باراسورامان وزيثامل وبيري، ١٩٨٨. الأداة. اختُزلت المحدّدات العشرة بالتحليل العاملي إلى خمسة أبعادٍ واثنتين وعشرين عبارةً مزدوجة، إحداهما تسأل ما ينبغي أن تقدّمه منشأةٌ ممتازة في الصناعة والأخرى تسأل ما قدّمته هذه المنشأة. ودرجة كل عبارةٍ هي الفرق، ودرجة البُعد هي المتوسّط.
كرونين وتايلور، ١٩٩٢. أهمّ تحدٍّ، وهو تجريبي لا مفاهيمي. يريان أن جودة الخدمة تُقاس أحسنَ بالأداء وحده، وهي أداةٌ يسمّيانها SERVPERF، ويبيّنان أن نسخة الأداء وحده تفسّر تباينًا في الجودة الإجمالية وفي نية الشراء أكثر مما تفسّره درجة الفرق. وحجتهما الإحصائية على درجات الفرق قوية: فطرحُ متغيّرين مقيسين يحمل خطأ كليهما ويكون عادةً أقلّ ثباتًا من أيٍّ منهما.
باراسورامان وزيثامل وبيري، ١٩٩٤. الردّ، ويسلّم بأرضٍ في الإحصاء ويدافع عن مكوّن التوقّعات لأسبابٍ تشخيصية: فالأداء وحده يخبر المدير بالدرجة، والفجوة تخبره بموضع النقص. وتُدخل الورقة كذلك منطقة تسامحٍ بين الخدمة المرغوبة والكافية، وهذا تحسينٌ حقيقي على معيار توقّعٍ واحد.
باتل، ١٩٩٦. النقد المجمَّع، وهو الاستشهاد المعياري عند الإقرار بمشكلات الأداة.
| Dimension | What it captures | Note |
|---|---|---|
| Reliability | Performing the promised service dependably and accurately | Consistently the most important dimension across studies, and the one customers weight most |
| Assurance | Knowledge and courtesy of staff and their ability to inspire trust | Dominant where the customer cannot judge the technical outcome, as in medicine, law and finance |
| Tangibles | Facilities, equipment, materials and staff appearance | The only dimension the customer can assess before purchase, so it carries the pre-purchase signal |
| Empathy | Individualized attention and understanding of the customer | Its cultural specificity is the instrument's weakest point in translation |
| Responsiveness | Willingness to help and to provide prompt service | Judged against expectations of speed that vary sharply by context |
Reliability is not one dimension among five. Across a large body of replications it is the strongest driver of overall quality judgments, and the practical implication is direct: a service that is warm, attractive and fast but unreliable scores badly, while a service that reliably does what it promised can compensate for weakness elsewhere. Doing the basic thing right, every time, outranks everything the instrument measures.
The expectation standard is the instrument's difficulty. What the expectation item is supposed to measure has never been settled. It has been read as what an excellent firm should offer, as what the customer predicts will happen, as what the customer considers adequate, and as what they desire. These are different standards and they produce different gaps, which is one reason the difference scores behave poorly. The 1994 zone of tolerance formulation, which distinguishes desired from adequate service and treats the space between as acceptable, is the more defensible version.
| البُعد | ما يلتقطه | ملاحظة |
|---|---|---|
| الاعتمادية | أداء الخدمة الموعودة أداءً موثوقًا دقيقًا | أهمّ الأبعاد على الدوام عبر الدراسات، وأثقلها وزنًا عند العملاء |
| الأمان | معرفة الموظفين ولطفهم وقدرتهم على بعث الثقة | غالبٌ حيث يعجز العميل عن الحكم على النتيجة التقنية، كالطبّ والقانون والمالية |
| الملموسات | المرافق والمعدّات والمواد ومظهر الموظفين | البُعد الوحيد الذي يستطيع العميل تقييمه قبل الشراء، فيحمل الإشارة قبل الشرائية |
| التعاطف | الاهتمام الفردي وفهم العميل | خصوصيته الثقافية أضعف مواضع الأداة في الترجمة |
| الاستجابة | الرغبة في المساعدة وتقديم خدمةٍ عاجلة | يُحكم عليه بتوقّعات سرعةٍ تتباين تباينًا حادًّا بالسياق |
والاعتمادية ليست بُعدًا من خمسة. ففي جسمٍ كبير من التكرارات هي أقوى محرّكات الحكم على الجودة الإجمالية، والدلالة العملية مباشرة: فخدمةٌ دافئة جذّابة سريعة غير موثوقة تنال درجةً سيئة، وخدمةٌ تفعل ما وعدت به بوثوقٍ تستطيع تعويض ضعفٍ في غير ذلك. فإتقان الأمر الأساسي، في كل مرة، يفوق كلَّ ما تقيسه الأداة.
ومعيار التوقّع هو عُسر الأداة. فما تقيسه عبارةُ التوقّع لم يُحسم قط. فقد قُرئت ما ينبغي أن تقدّمه منشأةٌ ممتازة، وما يتوقّع العميل حدوثه، وما يعدّه كافيًا، وما يرغبه. وهذه معايير مختلفة تُنتج فجواتٍ مختلفة، وهذا أحد أسباب سوء سلوك درجات الفرق. وصياغةُ منطقة التسامح سنة ١٩٩٤، وهي تفرّق بين الخدمة المرغوبة والكافية وتعدّ ما بينهما مقبولًا، هي النسخة الأولى بالدفاع.
Do not assume the five factor structure holds. The dimensionality has failed to replicate in many contexts, with items loading differently or collapsing into fewer factors. Running a confirmatory factor analysis and reporting the result is not a formality here, it is the first finding, and reporting a poor fit honestly is more valuable than forcing the published structure.
Weight the dimensions. The standard scoring averages the five, which assumes they matter equally. They do not, and reliability usually dominates. Asking respondents to allocate importance across the dimensions, or deriving weights from a regression on overall quality, produces a more informative result and a stronger paper.
Translation is a research act, not an administrative one. The empathy dimension in particular carries assumptions about individualized attention and appropriate personal distance that do not transfer uniformly, and expectations of responsiveness vary with local norms about time. Back translation is the minimum; establishing measurement invariance before comparing groups is what a good journal expects, and studies that compare SERVQUAL scores across countries without testing invariance are comparing scores that may not mean the same thing.
Where the contribution usually sits. Establishing the dimensional structure in a context where it has not been tested is publishable if done rigorously. Measurement invariance across cultures is under-supplied relative to how often cross-cultural comparisons are made. And digital and automated service delivery changes what the dimensions mean, since assurance conveyed by a person and assurance conveyed by an interface are not the same construct.
لا تفترض ثبات البنية الخماسية. فقد أخفقت الأبعادية في التكرار في سياقاتٍ كثيرة، بتحميل العبارات تحميلًا مختلفًا أو انطوائها في عواملَ أقلّ. وإجراءُ تحليلٍ عاملي توكيدي وذكرُ نتيجته ليس إجراءً شكليًّا هنا، بل هو النتيجة الأولى، وذكرُ سوء المطابقة بأمانةٍ أثمنُ من فرض البنية المنشورة.
وزِن الأبعاد. فالتسجيل المعياري يأخذ متوسّط الخمسة، وهذا يفترض تساويها في الأهمية. وهي غير متساوية، والاعتمادية تغلب عادةً. وسؤالُ المستجيبين توزيعَ الأهمية على الأبعاد، أو استخراجُ الأوزان من انحدارٍ على الجودة الإجمالية، يعطي نتيجةً أكثر إخبارًا وورقةً أقوى.
والترجمة فعلٌ بحثي لا فعلٌ إداري. فبُعد التعاطف خاصةً يحمل افتراضاتٍ في الاهتمام الفردي والمسافة الشخصية الملائمة لا تنتقل انتقالًا متجانسًا، وتوقّعات الاستجابة تتباين بأعراف الوقت المحلية. والترجمة العكسية هي الحدّ الأدنى؛ وإثباتُ ثبات القياس قبل مقارنة المجموعات هو ما تنتظره دوريةٌ جيدة، والدراسات التي تقارن درجات SERVQUAL بين البلدان بلا اختبار الثبات تقارن درجاتٍ قد لا تعني الشيء نفسه.
أين يقع الإسهام عادةً. إثبات البنية البُعدية في سياقٍ لم تُختبر فيه صالحٌ للنشر إن صُنع بدقة. وثبات القياس عبر الثقافات معروضٌ أقلّ مما تُجرى المقارنات عبر الثقافية. والتقديم الرقمي والآلي للخدمة يغيّر معنى الأبعاد، فالأمان الذي ينقله إنسان والأمان الذي تنقله واجهةٌ ليسا المفهوم نفسه.
Difference scores are psychometrically weak. This is the core technical objection. A difference of two measured variables inherits the error of both, usually has lower reliability than either component, and can produce a spurious negative correlation with the components. Cronin and Taylor's performance-only alternative typically outperforms it empirically, and this is not a matter of preference.
Expectations are measured after the service. In the standard administration both items are answered at the same time, after the encounter, so the expectation report is contaminated by the experience. Measuring expectations before and perceptions after is methodologically correct and almost never done because it is impractical.
The dimensions do not replicate. The five factor structure was derived in a small number of American service industries and has failed to reproduce in many others. Treating it as established rather than as a hypothesis is a widespread error.
It measures the process and not the outcome. SERVQUAL captures how the service was delivered. It says almost nothing about whether the technical outcome was correct, which is what actually matters in medicine, engineering, law and repair. Grönroos's distinction between functional and technical quality names what the instrument leaves out.
It assumes an encounter with a person. Automated, self-service and algorithmic delivery does not fit several items, and adapted instruments for electronic service exist for this reason.
درجات الفرق ضعيفة قياسيًّا. وهذا الاعتراض التقني الجوهري. فالفرق بين متغيّرين مقيسين يرث خطأ كليهما، ويكون ثباته عادةً أقلّ من ثبات أيٍّ من مكوّنيه، وقد يُنتج ارتباطًا سالبًا زائفًا مع المكوّنات. وبديلُ كرونين وتايلور القائم على الأداء وحده يتفوّق عليه تجريبيًّا عادةً، وهذه ليست مسألة تفضيل.
والتوقّعات تُقاس بعد الخدمة. ففي التطبيق المعياري تُجاب العبارتان في الوقت نفسه، بعد اللقاء، فيتلوّث تقرير التوقّع بالخبرة. وقياسُ التوقّعات قبلُ والإدراكات بعدُ صحيحٌ منهجيًّا ولا يكاد يُصنع لأنه غير عملي.
والأبعاد لا تتكرّر. فالبنية الخماسية استُخرجت في عددٍ قليل من صناعات الخدمة الأمريكية وأخفقت في إعادة إنتاجها في كثيرٍ غيرها. ومعاملتها راسخةً لا فرضيةً خطأٌ واسع الانتشار.
وتقيس العملية لا النتيجة. فSERVQUAL يلتقط كيف قُدّمت الخدمة. ولا يكاد يقول شيئًا عن صحة النتيجة التقنية، وهي المهمّة فعلًا في الطبّ والهندسة والقانون والإصلاح. وتفريقُ غرونروس بين الجودة الوظيفية والتقنية يسمّي ما تتركه الأداة.
وتفترض لقاءً بإنسان. فالتقديم الآلي والخدمة الذاتية والخوارزمية لا تلائم عدة عبارات، ولهذا وُجدت أدواتٌ مكيَّفة للخدمة الإلكترونية.