From AI-Generated Items to Measurement Validity Evidence: The Roles of Content Validity, Cognitive Interviewing, and Human Verification in Questionnaire Development
NGUYỄN VĂN HÙNG
HIGHLIGHTS
- Phân biệt AI-assisted item generation với measurement validation.
- Phát triển chuỗi bằng chứng CVQ → RPQ → PMQ cho bảng hỏi có AI hỗ trợ.
- Định vị Human Verification Quality như điều kiện biên của AI-assisted development.
- Phân biệt development efficiency với measurement-quality improvement.
- Đề xuất blinded source evaluation và traceable item provenance.
- Phát triển khung evidence-governed generative psychometrics dựa trên trách nhiệm khoa học của con người.
TÓM TẮT
Bối cảnh
Trí tuệ nhân tạo tạo sinh (Generative Artificial Intelligence – GenAI), đặc biệt là các mô hình ngôn ngữ lớn (large language models – LLMs), đang làm thay đổi nhanh quá trình phát triển bảng hỏi và thang đo. Từ một định nghĩa cấu trúc và một số ràng buộc thiết kế, GenAI có thể tạo hàng chục hoặc hàng trăm mục hỏi ứng viên, đề xuất cách diễn đạt thay thế, điều chỉnh độ khó đọc, hỗ trợ dịch thuật, phát hiện trùng lặp ngữ nghĩa, đề xuất response options và tham gia vào quá trình tinh chỉnh công cụ. Bằng chứng thực nghiệm gần đây cho thấy mục hỏi do AI tạo có thể đạt cấu trúc nhân tố, độ tin cậy và một số bằng chứng về internal structure cạnh tranh với các thang đo phát triển theo quy trình truyền thống. Tuy nhiên, năng lực tạo ngôn ngữ và kết quả tâm trắc thuận lợi không tự động tạo thành một lập luận giá trị đầy đủ cho các diễn giải và cách sử dụng điểm số dự kiến.
Khoảng trống nghiên cứu
Literature hiện tại chưa phân biệt đầy đủ bốn lớp vấn đề trong AI-assisted questionnaire development: nguồn tạo mục hỏi, chất lượng biểu diễn nội dung, chất lượng quá trình phản hồi của người trả lời và chất lượng bằng chứng tâm trắc. Đồng thời, khái niệm human-in-the-loop thường được sử dụng như một bảo đảm chung mà chưa xác định rõ con người đã thực hiện những hành vi kiểm chứng phương pháp nào. Vì vậy, khoảng trống quan trọng không còn nằm ở câu hỏi AI có khả năng tạo mục hỏi hay không, mà nằm ở việc thiếu một kiến trúc lý thuyết giải thích khi nào AI-assisted item generation tạo ra measurement value thực chất thay vì chỉ làm tăng tốc quá trình sản xuất nội dung đo lường.
Mục tiêu và khung lý thuyết
Bài viết phát triển Evidence-Governed Generative Psychometrics Framework, trong đó:
Item Development Condition (IDC)
→ Content Validity Quality (CVQ)
→ Response-Process Quality (RPQ)
→ Psychometric Measurement Quality (PMQ)
với Human Verification Quality (HVQ) là điều kiện biên quyết định liệu AI assistance trở thành quality-enhancing augmentation hay chỉ process acceleration. Development Efficiency (DE) được xem là outcome song song, có quan hệ với nhưng không đồng nhất với measurement quality. PMQ được khái niệm hóa như một multidimensional psychometric evidence profile, không mặc định là một reflective latent construct duy nhất.
Phương pháp đề xuất
Một multi-phase comparative instrument-development experiment được đề xuất, so sánh ba điều kiện: Human-Generated, AI-Generated và Human–AI Co-Designed. Quy trình gồm construct specification, randomized item development, blinded expert review, content-validity assessment, cognitive interviewing, evidence-based item revision, pilot administration và psychometric evaluation. Các quyết định liên quan AI được lưu theo chuỗi:
Prompt → AI Output → Human Decision → Rationale → Expert Evidence → Cognitive Evidence → Pilot Evidence → Revision → Final Item.
Các nguồn validity evidence được phân tích riêng biệt rồi tích hợp thành một validity argument thay vì cộng cơ học thành một chỉ số “validity”.
Đóng góp
Bài viết xác lập sáu phân biệt: AI Item Generation ≠ Instrument Validation; Linguistic Quality ≠ Measurement Quality; Expert Agreement ≠ Response-Process Validity Evidence; Good Model Fit ≠ Valid Measurement; Human Review ≠ Methodological Verification; Efficiency Gain ≠ Measurement Gain. Trên nền đó, bài viết phát triển Human Verification Quality, traceable item provenance và evidence governance như các cơ chế trung tâm của generative psychometrics có trách nhiệm. Luận đề cốt lõi là giá trị khoa học của GenAI trong questionnaire development không nên được đánh giá chủ yếu bằng số lượng mục hỏi được tạo, độ trôi chảy ngôn ngữ hoặc tốc độ hoàn thành thang đo, mà bằng khả năng đóng góp vào một chuỗi bằng chứng minh bạch, có thể truy nguyên, được con người kiểm chứng và đủ sức bảo vệ các diễn giải cũng như cách sử dụng điểm số đã xác định.
Từ khóa: Generative AI; questionnaire development; scale development; item generation; content validity; cognitive interviewing; response-process evidence; psychometrics; human verification; measurement validity evidence; generative psychometrics.

ABSTRACT
Background
Generative Artificial Intelligence (GenAI), particularly large language models (LLMs), is rapidly transforming questionnaire and scale development. Given a construct definition and a set of design constraints, GenAI can generate dozens or hundreds of candidate items, propose alternative wording, adjust readability, support translation, identify semantic redundancy, suggest response options, and assist with iterative instrument refinement. Emerging empirical evidence indicates that AI-generated items can exhibit factor structures, reliability, and evidence based on internal structure that are competitive with those of conventionally developed measures. However, linguistic generativity and favorable psychometric performance do not, by themselves, establish a comprehensive validity argument for the intended interpretations and uses of scores.
Research Gap
The existing literature has not sufficiently distinguished four analytically different components of AI-assisted questionnaire development: the source of item generation, the quality of content representation, the quality of respondents’ response processes, and the quality of psychometric evidence. Moreover, the concept of human-in-the-loop is frequently invoked as a general safeguard without specifying the methodological verification behaviors through which human involvement contributes to measurement quality. The critical research gap therefore no longer concerns whether AI can generate questionnaire items, but rather the absence of a theoretical architecture explaining when AI-assisted item generation creates substantive measurement value rather than merely accelerating the production of measurement content.
Aim and Theoretical Framework
This article develops an Evidence-Governed Generative Psychometrics Framework comprising the following sequence:
Item Development Condition (IDC)
→ Content Validity Quality (CVQ)
→ Response-Process Quality (RPQ)
→ Psychometric Measurement Quality (PMQ)
Within this framework, Human Verification Quality (HVQ) is conceptualized as a boundary condition determining whether AI assistance functions as quality-enhancing augmentation or merely as process acceleration. Development Efficiency (DE) is conceptualized as a parallel outcome that is related to, but analytically distinct from, measurement quality. PMQ is defined as a multidimensional psychometric evidence profile rather than being assumed to constitute a single reflective latent construct.
Proposed Method
A multi-phase comparative instrument-development experiment is proposed to compare three conditions: Human-Generated, AI-Generated, and Human–AI Co-Designed. The research process comprises construct specification, randomized item development, blinded expert review, content-validity assessment, cognitive interviewing, evidence-based item revision, pilot administration, and psychometric evaluation. AI-related decisions are documented through a traceable sequence:
Prompt → AI Output → Human Decision → Rationale → Expert Evidence → Cognitive Evidence → Pilot Evidence → Revision → Final Item.
Distinct sources of validity evidence are evaluated separately and subsequently integrated into an overall validity argument rather than being mechanically aggregated into a single “validity” index.
Contribution
The article advances six methodological distinctions: AI Item Generation ≠ Instrument Validation; Linguistic Quality ≠ Measurement Quality; Expert Agreement ≠ Response-Process Validity Evidence; Good Model Fit ≠ Valid Measurement; Human Review ≠ Methodological Verification; and Efficiency Gain ≠ Measurement Gain. Building on these distinctions, the framework develops Human Verification Quality, traceable item provenance, and evidence governance as central mechanisms of responsible generative psychometrics. The central argument is that the scientific value of GenAI in questionnaire development should not be judged primarily by the number of items generated, their linguistic fluency, or the speed with which a scale can be assembled. Rather, it should be judged by the extent to which AI-assisted development contributes to a transparent, traceable, human-verified, and empirically defensible chain of validity evidence supporting the intended interpretations and uses of scores.
Keywords: generative artificial intelligence; questionnaire development; scale development; item generation; content validity; cognitive interviewing; response-process evidence; psychometrics; human verification; measurement validity evidence; generative psychometrics.
- GIỚI THIỆU
1.1. Từ viết mục hỏi thủ công đến generative psychometrics
Phát triển bảng hỏi có vẻ là một hoạt động ngôn ngữ tương đối đơn giản: xác định điều cần đo, viết một tập câu hỏi và yêu cầu người tham gia phản hồi. Tuy nhiên, measurement science cho thấy một mục hỏi tốt không chỉ cần “đọc hay” hay “dễ hiểu”. Một công cụ đo lường có khả năng hỗ trợ suy luận khoa học phải kết nối chặt chẽ ít nhất năm lớp quyết định: construct definition, content representation, respondent cognition, psychometric structure và score interpretation. Vì thế, scale development được xem là một quá trình tuần tự, lặp lại và tích lũy bằng chứng hơn là một nhiệm vụ tạo câu chữ đơn thuần (Boateng et al., 2018; Clark & Watson, 2019; DeVellis & Thorpe, 2022).
GenAI thay đổi đáng kể lớp đầu của quy trình này. Các LLMs có thể chuyển một construct definition ngắn thành một initial item pool trong thời gian rất ngắn; tạo thêm item theo từng facet; cung cấp semantic alternatives; rút gọn wording; điều chỉnh reading level; phát hiện apparent redundancy; hỗ trợ translation; đề xuất response categories; và đóng vai một preliminary reviewer để nhận diện double-barreled wording hoặc ambiguity.
Hoffmann et al. (2024) cho thấy ChatGPT có thể được tích hợp vào một quy trình phát triển thang đo mới và công cụ tạo ra có thể tiếp tục được kiểm tra bằng các nghiên cứu thực nghiệm. Salah et al. (2025) sử dụng thiết kế mixed-method để đánh giá khả năng GPT-4 tạo và thích nghi thang đo, cung cấp bằng chứng cho thấy AI-generated items có thể đạt một số đặc tính đo lường thuận lợi, đồng thời nhấn mạnh vai trò không thể bỏ qua của human expertise. Terry et al. (2025) tiếp tục cho thấy các mục hỏi do AI tạo có thể đạt kết quả cạnh tranh với gold-standard measures trong một số điều kiện. Đặc biệt, Russell-Lasalandra et al. (2026) phát triển AI-GENIE, kết hợp LLMs, text embeddings và network psychometrics để tạo, sàng lọc và đánh giá item pools; nghiên cứu của họ cho thấy những thang đo được phát triển qua pipeline này có thể đạt bằng chứng về internal structure cạnh tranh với các measures do chuyên gia xây dựng.
Các phát hiện đó có ý nghĩa lớn: AI-generated items không thể bị bác bỏ chỉ vì nguồn tạo của chúng là máy. Tuy nhiên, kết luận ngược lại—rằng items có chất lượng ngôn ngữ hoặc internal structure tốt thì mặc nhiên đã “validated”—cũng không được hỗ trợ.
Từ đó hình thành phát biểu học thuật tích hợp thứ nhất:
GenAI có thể làm giảm mạnh chi phí sản xuất measurement content, nhưng khả năng tạo nội dung không tự động làm giảm yêu cầu về validity evidence.
Điều này chuyển câu hỏi nghiên cứu từ:
“AI có thể viết mục hỏi hay không?”
sang:
“Một mục hỏi có AI tham gia tạo phải vượt qua những lớp bằng chứng nào trước khi được dùng để hỗ trợ một diễn giải đo lường có thể bảo vệ?”
Phân biệt đầu tiên vì vậy là:
AI Item Generation ≠ Instrument Validation.
1.2. Nghịch lý efficiency–validity
Sự tham gia của GenAI tạo ra một nghịch lý phương pháp.
Một mặt, scale development truyền thống tiêu tốn nhiều thời gian cho việc tạo candidate items, rà soát wording, tìm semantic alternatives, cân bằng các facets và loại bỏ nội dung trùng lặp. GenAI có thể giảm đáng kể chi phí của những hoạt động này. Kuru (2025) cho thấy AI có khả năng được tích hợp xuyên suốt survey-development cycle, trong khi Valenzuela et al. (2025) phân tích nhiều cơ hội sử dụng LLMs trong survey research nhưng đồng thời cảnh báo về validity, transparency và reproducibility.
Mặt khác, khả năng tạo ngôn ngữ trôi chảy của LLMs có thể tạo ra một dạng linguistic plausibility. Một item có thể đọc tự nhiên, cân đối và chuyên nghiệp nhưng vẫn đại diện không đầy đủ cho construct. Hai items có thể sử dụng wording khác nhau nhưng đo cùng một micro-behavior. Một item có thể hoàn toàn rõ đối với chuyên gia nhưng kích hoạt cách hiểu khác ở target respondents. Một item pool có thể cho reliability cao do semantic homogeneity nhưng lại bỏ sót những facets quan trọng.
Từ đó xuất hiện phát biểu tích hợp thứ hai:
Linguistic competence của AI và measurement quality thuộc hai cấp bằng chứng khác nhau.
Do vậy:
Linguistic Quality ≠ Measurement Quality.
Một technology có thể rất giỏi tạo text mà chưa chắc rất giỏi bảo đảm construct representation.
1.3. Khoảng trống generation–validation
Trong scale-development methodology, item generation chỉ là một giai đoạn. Sau đó thường cần expert review, pretesting, target-population evaluation và psychometric evaluation. Nhưng GenAI làm item generation nhanh đến mức researcher có thể bị cám dỗ rút ngắn những bước chậm hơn phía sau.
Một workflow rủi ro là:
Construct Prompt
→ AI-Generated Items
→ Survey Administration
→ CFA
→ “Validated Scale”.
Workflow này nhảy trực tiếp từ production sang internal structure.
Nó không trả lời các câu hỏi nền tảng: construct domain có được prespecify đầy đủ không? AI có bao phủ các facets khó diễn đạt không? Có construct contamination không? Target respondents có hiểu items như intended không? Model fit tốt có thể phản ánh substantive dimensionality hay chỉ semantic redundancy?
Clark và Watson (2019) nhấn mạnh mối liên hệ giữa construct conceptualization, item content và empirical structure. DeVellis và Thorpe (2022) cũng xem item generation là một phần của chu trình phát triển công cụ rộng hơn. Vì vậy, GenAI không làm validation trở nên lỗi thời.
Ngược lại, nó làm validation quan trọng hơn.
Đây là phát biểu tích hợp thứ ba:
Khi item generation trở nên rẻ, nhanh và gần như không giới hạn, bottleneck của measurement science chuyển từ production sang verification.
1.4. Khoảng trống content–response process
Content validity và response-process evidence thường xuất hiện gần nhau trong development workflow nhưng không trả lời cùng một câu hỏi.
Expert review hỏi:
“Theo theory và domain knowledge, item này có đại diện construct hay không?”
Cognitive interviewing hỏi:
“Khi target respondent đọc item, họ thực sự hiểu gì, truy hồi thông tin gì, hình thành phán đoán nào và ánh xạ phán đoán đó vào response category ra sao?”
Một expert về AI literacy có thể hiểu “algorithmic bias” theo nghĩa kỹ thuật. Một sinh viên năm nhất có thể hiểu cùng thuật ngữ như “AI thiên vị cá nhân tôi”. Expert agreement cao do đó không chứng minh respondent cognition phù hợp.
Từ đây hình thành phát biểu tích hợp thứ tư:
Validity evidence phải xét cả intended representation của developer và enacted interpretation của respondent.
Do vậy:
Expert Agreement ≠ Response-Process Validity Evidence.
1.5. Khoảng trống psychometric-fit
Một lỗi suy luận phổ biến là biến một tập psychometric thresholds thành chứng minh cuối cùng của validity. Researchers có thể kết luận “scale is valid” sau khi quan sát CFI hoặc TLI thuận lợi, RMSEA thấp, alpha hoặc omega cao, AVE vượt một ngưỡng hoặc HTMT dưới một cut-off.
Các indicators đó hữu ích, nhưng mỗi indicator trả lời một nhóm câu hỏi hẹp.
Model fit không chứng minh construct domain được bao phủ đầy đủ.
Reliability không chứng minh item content đúng.
AVE không chứng minh respondents hiểu item theo intended meaning.
HTMT không sửa được conceptual overlap giữa hai constructs.
Standards for Educational and Psychological Testing định vị validity trong quan hệ giữa evidence, theory và intended score interpretations and uses (AERA et al., 2014).
Từ đó xuất hiện phát biểu tích hợp thứ năm:
Internal-structure evidence là một thành phần của validity argument, không phải toàn bộ validity argument.
Do đó:
Good Model Fit ≠ Valid Measurement.
1.6. Khoảng trống human verification
“Human-in-the-loop” ngày càng phổ biến trong discourse về AI governance. Tuy nhiên, sự hiện diện của con người không mặc nhiên đồng nghĩa kiểm chứng.
Một researcher có thể đọc 50 items do AI tạo, giữ lại 20 câu “nghe tốt” và mô tả quy trình là human oversight. Nếu researcher không kiểm tra construct mapping, domain coverage, semantic ambiguity, target-population interpretation hoặc empirical evidence, hoạt động trên gần với human selection hơn methodological verification.
Từ đây hình thành phát biểu tích hợp thứ sáu:
Human presence chỉ trở thành safeguard khoa học khi được chuyển hóa thành những hành vi kiểm chứng có bằng chứng, có lý do và có thể truy nguyên.
Do đó:
Human Review ≠ Methodological Verification.
1.7. Khoảng trống efficiency–quality
AI có lợi thế rõ ràng về raw generation speed. Nhưng development speed và measurement quality không phải cùng một outcome.
Một item pool có thể được tạo trong 30 giây nhưng cần nhiều giờ sửa construct drift, redundancy và ambiguity. Ngược lại, một workflow Human–AI có thể giảm person-hours mà vẫn duy trì hoặc nâng chất lượng validity evidence.
Do đó xuất hiện phân biệt thứ sáu:
Efficiency Gain ≠ Measurement Gain.
Đây không phải một chi tiết vận hành. Nó ngăn một productivity fallacy: nhanh hơn không mặc nhiên đồng nghĩa khoa học hơn.
1.8. Câu hỏi trung tâm
Bài viết đặt câu hỏi:
Trong những điều kiện nào AI-assisted item generation đóng góp vào measurement validity evidence thay vì chỉ làm tăng tốc questionnaire development?
Các câu hỏi cụ thể là:
RQ1. Human-generated, AI-generated và Human–AI co-designed item pools khác nhau như thế nào về Content Validity Quality?
RQ2. Các điều kiện item development khác nhau như thế nào về Response-Process Quality?
RQ3. Content Validity Quality liên quan như thế nào đến Response-Process Quality?
RQ4. Response-Process Quality liên quan như thế nào đến downstream Psychometric Measurement Quality?
RQ5. Human Verification Quality làm thay đổi mối quan hệ giữa AI involvement và downstream measurement quality như thế nào?
RQ6. AI tạo ra development-efficiency gains, measurement-quality gains hay cả hai?
- KIẾN TRÚC LẬP LUẬN TÍCH HỢP BẢY BƯỚC
Bài viết không xem các nguồn như những câu trích dẫn đứng độc lập. Lý thuyết, empirical evidence và methodological principles được tổ chức thành một coordinated argument architecture gồm bảy bước.
2.1. Bước 1 – Xác lập năng lực công nghệ
Bước đầu tiên thừa nhận một thực tế mà literature mới đã chỉ ra: AI có thể tạo measurement content có chất lượng đáng kể.
Hoffmann et al. (2024), Salah et al. (2025), Terry et al. (2025) và Russell-Lasalandra et al. (2026) đều cung cấp evidence chống lại giả định đơn giản rằng AI-generated items tất yếu inferior.
Luận điểm này quan trọng vì một framework tốt không nên được xây dựng trên AI aversion.
Nguồn tạo không phải quality verdict.
Source Identity ≠ Measurement Quality.
2.2. Bước 2 – Xác định giới hạn của generativity
LLMs được tối ưu để sinh chuỗi ngôn ngữ phù hợp với context. Measurement science cần hơn thế: construct representation, interpretability và evidential defensibility.
Từ đó hình thành phát biểu tích hợp thứ bảy:
Generative capability mở rộng không gian candidate items nhưng không tự động xác định epistemic status của các items đó.
Một câu trôi chảy mới chỉ là candidate content.
2.3. Bước 3 – Chuyển từ sản phẩm sang validity argument
Messick (1989) và Standards (AERA et al., 2014) làm rõ rằng validity không nên được hiểu như một nhãn cố định gắn vào instrument. Điều quan trọng là evidence và theory hỗ trợ intended interpretations and uses.
Từ đó xuất hiện phát biểu tích hợp thứ tám:
Một scale không được “validated” một lần cho mọi population, context và purpose; validity phải được lập luận theo specific interpretation and use.
Do vậy, bài viết ưu tiên thuật ngữ Measurement Validity Evidence hơn cách hiểu nhị phân “valid/invalid instrument”.
2.4. Bước 4 – Thiết lập progression từ content tới cognition và psychometrics
Một measurement argument mạnh cần kết nối:
Content Representation
→ Respondent Cognition
→ Empirical Response Structure.
Content evidence cho biết item nên đại diện điều gì.
Response-process evidence cho biết người tham gia thực sự làm gì khi trả lời.
Psychometric evidence cho biết phản hồi tạo ra statistical patterns nào.
Đây là phát biểu tích hợp thứ chín:
Psychometric evidence có sức thuyết phục lớn nhất khi nó là downstream evidence của item content đã được biện minh và respondent cognition đã được kiểm tra.
2.5. Bước 5 – Chuyển human oversight thành verification mechanism
“Human-in-the-loop” là quá rộng để trở thành một construct phân tích.
Human Verification Quality yêu cầu observable behaviors:
construct checking;
domain-gap detection;
alternative comparison;
output rejection;
evidence-linked revision;
uncertainty documentation;
decision traceability.
Đây là phát biểu tích hợp thứ mười:
Human oversight tạo giá trị khoa học khi nó tạo independent evidential checks, không phải khi con người chỉ xác nhận output của AI.
2.6. Bước 6 – Tách efficiency khỏi measurement quality
Development efficiency phải tính cả generation, verification, expert review, cognitive interviewing, revision và rework.
Từ đó hình thành phát biểu tích hợp thứ mười một:
Technological productivity và epistemic quality là hai outcome domains riêng biệt.
Một workflow tốt nhất không nhất thiết là workflow nhanh nhất.
2.7. Bước 7 – Đặt scientific accountability ở cuối chuỗi
Nếu AI tham gia tạo research instrument, AI trở thành một phần của research process. Tuy nhiên, AI không chịu trách nhiệm về construct definition, item retention, validity claims hoặc consequences of score interpretation.
Từ đó hình thành phát biểu tích hợp thứ mười hai:
Automation of Generation ≠ Delegation of Scientific Accountability.
Và phát biểu thứ mười ba tổng hợp toàn bài:
Giá trị khoa học của generative psychometrics không nằm ở số lượng item AI tạo được mà nằm ở việc hệ thống Người–AI có tạo được một chuỗi bằng chứng minh bạch, có thể truy nguyên và đủ mạnh để bảo vệ intended score interpretations and uses hay không.

- VALIDITY NHƯ MỘT LẬP LUẬN BẰNG CHỨNG
3.1. Từ “valid instrument” đến validity evidence
Cách nói “the scale was validated” thường tạo cảm giác validity là một thuộc tính cố định của thang đo.
Cách hiểu phù hợp hơn là validity liên quan đến strength của evidence và theory hỗ trợ interpretations of scores trong một use cụ thể.
Điều này đặc biệt quan trọng với AI-generated instruments. Một scale hoạt động tốt trong một sample không cho phép suy rộng rằng AI-generated scale đó valid ở mọi population.
Kết luận thích hợp hơn là:
Trong population, context, language và intended use được nghiên cứu, những nguồn evidence cụ thể hỗ trợ hoặc chưa hỗ trợ score interpretation ở mức độ nào?
3.2. Nhiều nguồn validity evidence
Đối với questionnaire development, ít nhất năm nhóm evidence có thể quan trọng.
Evidence based on content kiểm tra relevance, representativeness, coverage và construct alignment.
Evidence based on response processes kiểm tra cách respondents hiểu, truy hồi và hình thành phán đoán.
Evidence based on internal structure kiểm tra dimensionality và quan hệ giữa items.
Evidence based on relations to other variables kiểm tra các theoretical predictions bên ngoài instrument.
Evidence liên quan consequences trở nên quan trọng khi scores được dùng trong classification hoặc consequential decisions.
Không phải study nào cũng cần đồng thời thu thập mọi nguồn bằng chứng. Nhưng một nguồn evidence không được trình bày như toàn bộ validity argument.
3.3. Reliability và validity
Reliability phản ánh consistency hoặc precision.
Một AI system rất giỏi paraphrase có thể tạo năm item gần như cùng nghĩa. Các items đó có thể covary mạnh và tạo reliability cao.
Tuy nhiên, chúng vẫn chỉ bao phủ một vùng rất hẹp của construct.
McNeish (2018) phê bình việc dùng Cronbach’s alpha như một thủ tục mặc định, đặc biệt khi assumptions của coefficient không phù hợp measurement model.
Do đó:
Reliability ≠ Validity.
Trong empirical phase, omega hoặc model-based reliability nên được xem xét khi phù hợp, trong khi alpha có thể được báo cáo phục vụ comparability nếu target literature thường sử dụng.
- AI-ASSISTED ITEM GENERATION
4.1. Những năng lực có ích
GenAI có thể hỗ trợ ít nhất chín nhiệm vụ:
- tạo initial item pools;
- tạo item riêng cho từng facet;
- đề xuất alternative wording;
- điều chỉnh reading level;
- phát hiện apparent redundancy;
- hỗ trợ translation/adaptation;
- đề xuất response options;
- nhận diện double-barreling hoặc ambiguity;
- đóng vai adversarial reviewer.
Điều cần tránh là chuyển capability thành validity claim.
Cách diễn đạt phù hợp nhất là:
AI-generated items constitute candidate measurement content whose evidential status remains to be established.
4.2. Construct underrepresentation
LLMs tạo output từ statistical patterns trong language data. Những manifestations thường xuyên xuất hiện trong text có thể được tạo dễ hơn các manifestations ngầm, phức tạp hoặc phụ thuộc bối cảnh.
Ví dụ, một prompt về AI literacy có thể tạo rất nhiều item liên quan đến “biết sử dụng AI”, “biết viết prompt” hoặc “biết AI có thể sai” nhưng ít item về data provenance, uncertainty evaluation, verification behavior hoặc accountability.
Một item pool lớn vì thế không đảm bảo construct breadth.
Quantity ≠ Domain Coverage.
4.3. Construct contamination
Prompt có thể vô tình pha trộn adjacent constructs.
Nếu researcher định nghĩa research self-efficacy bằng “confidence, competence and likelihood of success”, AI có thể sinh content pha lẫn self-efficacy, actual competence và outcome expectancy.
Construct specification phải xảy ra trước AI generation.
Một design yếu là để cùng AI:
định nghĩa construct;
tạo items;
đánh giá items;
và kết luận items phù hợp.
Đó không phải independent evidence.
4.4. Semantic redundancy và synthetic homogeneity
LLMs rất giỏi tạo paraphrases.
Ví dụ:
“I verify AI-generated information before using it.”
“I check AI-generated information before relying on it.”
“I confirm the accuracy of AI-generated information before using it.”
Ba câu nhìn khác nhau nhưng có thể đo cùng một micro-behavior.
Nếu giữ nhiều paraphrases, scale có thể đạt reliability cao do synthetic homogeneity.
Do đó:
Semantic Diversity ≠ Construct Breadth.
Russell-Lasalandra et al. (2026) đưa redundancy detection thành thành phần chính thức của AI-GENIE, cho thấy redundancy cần được xử lý bằng systematic procedure thay vì trực giác.
4.5. Cultural flattening
Một LLM có thể tạo generic items dễ áp dụng ở nhiều contexts. Tuy nhiên, measurement đôi khi cần contextual specificity.
Ví dụ, career self-reliance ở sinh viên Việt Nam có thể liên quan tới family expectations, internship opportunities, local labor-market uncertainty, digital transformation và cấu trúc chuyển tiếp từ university sang employment.
Generic wording có thể “correct” về grammar nhưng poor về contextual representation.
Target-population evaluation vì vậy vẫn cần thiết.
- CONTENT VALIDITY QUALITY
5.1. Định nghĩa
Content Validity Quality (CVQ) được định nghĩa là mức độ candidate item pool đại diện thích đáng, đầy đủ và có thể biện minh cho construct domain đã được prespecify trong target context.
CVQ không đồng nhất với CVI.
CVI là một indicator hữu ích của expert agreement. CVQ là construct rộng hơn về content representation.
5.2. Các miền CVQ
CVQ gồm ít nhất sáu miền.
Relevance: item có liên hệ trực tiếp với construct/facet?
Representativeness: item có phải manifestation tiêu biểu?
Coverage: item pool có bao phủ đầy đủ domain?
Construct alignment: item có tránh adjacent constructs?
Clarity: wording có rõ và đơn nghĩa?
Non-redundancy: mỗi item có đóng góp nội dung riêng?
5.3. Expert-panel composition
Panel nên kết hợp:
domain experts;
measurement/psychometric experts;
questionnaire/survey experts;
và linguistic/cultural experts khi design đa ngôn ngữ hoặc đa văn hóa.
Một panel chỉ gồm experts trong cùng discipline có thể đạt agreement rất cao nhưng bỏ qua survey-design problems.
5.4. Blinded source evaluation
Một cải tiến thiết kế quan trọng là không cho experts biết item thuộc điều kiện:
Human;
AI;
hay Human–AI.
Blinding giúp hạn chế hai bias đối nghịch:
AI enthusiasm bias;
AI aversion bias.
Nhờ đó, đối tượng được đánh giá là item content chứ không phải source identity.
5.5. CVI và threshold problem
I-CVI và S-CVI/Ave có thể mô tả expert agreement về relevance.
Tuy nhiên, không nên sử dụng threshold như pass/fail mechanism duy nhất.
Một item có CVI cao nhưng hoàn toàn redundant có thể không đáng giữ.
Một item có disagreement tương đối lớn nhưng đại diện một facet hiếm và quan trọng có thể vẫn cần tiếp tục xem xét.
Do vậy:
High CVI ≠ Complete Validity Evidence.
Qualitative rationales phải được giữ lại trong evidence trail.
6.COGNITIVE INTERVIEWING VÀ RESPONSE-PROCESS QUALITY
6.1. Vì sao cần cognitive interviewing?
Expert review nhìn item từ perspective của theory và measurement design.
Nhưng respondents mới là người tạo observed responses.
Willis (2005) cho thấy cognitive interviewing có thể hỗ trợ researcher hiểu cách người tham gia xử lý survey questions. MacDermid (2021) cũng nhấn mạnh giá trị bổ sung của cognitive interviewing đối với content-validation methods.
6.2. Response-process architecture
Quá trình phản hồi có thể được mô tả:
Comprehension → Retrieval → Judgment → Response Mapping.
Comprehension: respondent hiểu câu hỏi thế nào?
Retrieval: họ truy hồi trải nghiệm nào?
Judgment: họ tổng hợp thông tin ra sao?
Response mapping: họ chuyển judgment thành response category thế nào?
AI-generated wording có thể thất bại ở bất kỳ bước nào.
6.3. Clarity giả
Giả sử AI tạo item:
“Tôi chủ động xác minh tính đáng tin cậy của nội dung do AI tạo trước khi tích hợp vào công việc học tập.”
Câu có vẻ rõ.
Nhưng cognitive interviewing có thể phát hiện:
“xác minh” được hiểu là đọc lại;
“đáng tin cậy” được hiểu là nghe hợp lý;
“tích hợp” được hiểu là copy vào assignment.
Khi đó linguistic clarity không tương đương intended response process.
6.4. Định nghĩa RPQ
Response-Process Quality (RPQ) là mức độ cognitive processes thực tế của target respondents phù hợp với cognitive processes mà intended score interpretation giả định.
Các miền RPQ có thể gồm:
RPQ1 – Intended comprehension
RPQ2 – Relevant retrieval
RPQ3 – Construct-relevant judgment
RPQ4 – Appropriate response mapping
RPQ5 – Ambiguity control
RPQ6 – Contextual interpretability
RPQ không nên chỉ được đo bằng câu hỏi “item có dễ hiểu không?” trên Likert scale.
Nó cần coded qualitative evidence.
6.5. Cognitive-interview protocol
Protocol có thể kết hợp:
think-aloud;
comprehension probe;
paraphrasing;
retrieval probe;
judgment probe;
response-category probe;
confidence probe;
retrospective explanation.
Interviewers nên sử dụng standardized core probes nhưng vẫn cho phép emergent probing.
6.6. Sampling
Cognitive interviewing không nhắm tới statistical representativeness.
Mục tiêu là interpretive variation.
Sampling nên xem xét characteristics có thể làm thay đổi interpretation:
study level;
discipline;
language proficiency;
experience with construct;
cultural background;
digital literacy khi relevant.
Không có một con số interviews cố định phù hợp cho mọi nghiên cứu.
- PSYCHOMETRIC MEASUREMENT QUALITY
7.1. PMQ là một evidence profile
Psychometric Measurement Quality (PMQ) được định nghĩa là một multidimensional evidence profile phản ánh item functioning, dimensional structure, precision/consistency và theoretically relevant relations sau khi content và response-process issues đã được xử lý.
Điểm này cần được nhấn mạnh:
PMQ không mặc định là một reflective latent variable.
CFI, RMSEA, omega, AVE, HTMT và IRT parameters không phải những manifestations đồng nhất của một latent factor “quality”.
Do đó, empirical study nên báo PMQ như một profile hoặc set of outcomes được prespecify.
7.2. EFA, CFA và ESEM
EFA phù hợp khi structure chưa chắc chắn.
CFA phù hợp khi measurement model được prespecify.
ESEM có thể hữu ích khi theoretical model cho phép substantively meaningful cross-loadings.
Nguyên tắc:
Exploration ≠ Confirmation.
Nếu researcher chạy EFA, loại items, chỉnh model và sau đó chạy CFA trên cùng sample, CFA không phải independent confirmation.
Split sample hoặc independent validation sample nên được cân nhắc khi khả thi.
7.3. Item retention
Factor loading không nên là criterion duy nhất.
Item retention cần cân nhắc:
loading magnitude;
uncertainty;
cross-loading;
residual relationships;
theoretical importance;
facet coverage;
redundancy.
Một item loading .55 đại diện một facet độc đáo có thể quan trọng hơn item loading .85 nhưng lặp nội dung.
Psychometric Optimization ≠ Construct Optimization.
7.4. Reliability
Nghiên cứu nên tránh logic:
α > .70 → reliable → valid.
Omega hoặc model-based reliability có thể phù hợp hơn trong nhiều measurement settings, nhưng không coefficient nào tự nó xác lập validity.
Synthetic homogeneity do AI tạo đặc biệt làm distinction này quan trọng.
7.5. Convergent và discriminant evidence
CR và AVE có thể hữu ích trong những SEM applications phù hợp.
HTMT có thể hỗ trợ đánh giá discriminant relations.
Nhưng:
HTMT Threshold ≠ Construct Validity.
Nếu theoretical definitions overlap, statistical separation không tự động giải quyết conceptual ambiguity.
7.6. Measurement invariance
Nếu instrument được dùng giữa groups, invariance cần được xem xét:
configural;
metric;
scalar;
và residual khi substantive purpose đòi hỏi.
Một AI-generated questionnaire có thể đạt pooled model fit thuận lợi nhưng hoạt động khác nhau giữa linguistic hoặc cultural groups.
- HUMAN VERIFICATION QUALITY
8.1. Định nghĩa
Human Verification Quality (HVQ) là mức độ researcher chủ động kiểm tra, thách thức, so sánh, bác bỏ, sửa đổi và biện minh AI-generated measurement content bằng construct theory, target-population evidence và empirical evidence.
HVQ mạnh hơn generic “human review”.
Human reviewer có thể đọc.
Human verifier phải kiểm chứng.
8.2. Các miền HVQ
HVQ1 – Construct Verification: item có map đúng construct/facet?
HVQ2 – Content-Coverage Verification: AI có bỏ sót critical facet?
HVQ3 – Semantic Verification: item có ambiguity, hidden assumptions hoặc contamination?
HVQ4 – Population Verification: target population có hiểu item như intended?
HVQ5 – Evidence-Based Revision: item revisions có dựa trên evidence?
HVQ6 – Final Measurement Accountability: human team có giải thích được retention/revision/removal decisions?
8.3. Verification trace và item provenance
Mỗi AI-assisted item nên có lineage:
PROMPT
→ AI OUTPUT
→ HUMAN DECISION
→ RATIONALE
→ EXPERT EVIDENCE
→ COGNITIVE EVIDENCE
→ PILOT EVIDENCE
→ REVISION
→ FINAL ITEM.
Chuỗi này tạo traceable item provenance.
Do đó:
AI Suggested ≠ Human Accepted ≠ Methodologically Verified.
8.4. AI-use frequency không phải HVQ
Một researcher có thể sử dụng AI 100 lần và reject phần lớn suggestions sau verification.
Một researcher khác chỉ dùng AI một lần nhưng chấp nhận nguyên item pool.
Frequency không phản ánh verification quality.
HVQ phải ưu tiên process evidence.
8.5. Operationalization
HVQ indicators có thể gồm:
proportion of AI suggestions explicitly checked;
rejection with rationale;
construct-map checks;
alternative wording comparisons;
evidence-linked revisions;
documentation of unresolved uncertainty;
traceability completeness.
Decision logs có thể được blinded raters chấm bằng rubric.
- MÔ HÌNH KHÁI NIỆM VÀ RANH GIỚI CẤU TRÚC
9.1. Mô hình
Mô hình đề xuất:
ITEM DEVELOPMENT CONDITION – IDC
Human-Generated
vs. AI-Generated
vs. Human–AI Co-Designed
↓
CONTENT VALIDITY QUALITY – CVQ
↓
RESPONSE-PROCESS QUALITY – RPQ
↓
PSYCHOMETRIC MEASUREMENT QUALITY – PMQ
Song song:
AI Involvement → DEVELOPMENT EFFICIENCY – DE
Và:
HUMAN VERIFICATION QUALITY – HVQ
là boundary condition của các AI-involved pathways.
9.2. Ranh giới constructs
IDC là experimental condition, không phải psychological construct.
CVQ phản ánh chất lượng representational content, không đồng nhất với CVI.
RPQ phản ánh respondent cognition, không đồng nhất với expert judgment.
PMQ là multidimensional psychometric evidence profile, không phải một validity score đơn nhất.
HVQ phản ánh verification process của human developers, không phải AI-use frequency hoặc perceived oversight.
DE phản ánh resource efficiency, không phải measurement quality.
Việc xác lập ranh giới này giúp tránh construct overlap và circular reasoning.

- PHÁT TRIỂN GIẢ THUYẾT
10.1. IDC và CVQ
Human developers có thể có lợi thế về contextual knowledge.
AI có lợi thế về speed, breadth và linguistic alternatives.
Human–AI co-design có khả năng kết hợp hai nguồn competence, nhưng lợi thế đó phụ thuộc role configuration và verification.
Literature hiện tại chưa đủ để giả định universal superiority.
H1. Content Validity Quality differs across human-generated, AI-generated, and Human–AI co-designed item-development conditions.
10.2. CVQ và RPQ
Items có content alignment và clarity tốt về lý thuyết có khả năng tạo intended comprehension cao hơn.
Tuy nhiên, association không hoàn hảo vì expert cognition và respondent cognition khác nhau.
H2. Content Validity Quality is positively associated with Response-Process Quality.
10.3. RPQ và PMQ
Nếu respondents hiểu item theo intended meaning và sử dụng construct-relevant judgment, response covariance có khả năng phản ánh intended structure tốt hơn.
Systematic misunderstandings có thể tạo secondary dimensions hoặc residual relations.
H3. Response-Process Quality is positively associated with favorable downstream psychometric measurement evidence.
10.4. CVQ và PMQ
Content coverage, redundancy và construct purity có thể ảnh hưởng trực tiếp đến psychometric structure ngoài phần tác động thông qua RPQ.
H4. Content Validity Quality is positively associated with favorable downstream psychometric measurement evidence.
10.5. Mediation
Một phần quan hệ CVQ–PMQ có thể được giải thích bằng respondent cognition.
H5. Response-Process Quality mediates the association between Content Validity Quality and downstream psychometric measurement evidence.
H5 chỉ được estimate nếu study tạo đủ independent development units và temporal ordering thích hợp.
Nếu chỉ có một scale cuối cho mỗi condition, H5 phải được xem như conceptual proposition.
10.6. HVQ như boundary condition
AI involvement không được kỳ vọng có effect đồng nhất.
Verification thấp có thể làm construct drift hoặc redundancy không được phát hiện.
Verification cao có thể chuyển generative capability thành augmentation.
H6. Human Verification Quality moderates the relationship between AI involvement and downstream measurement quality, such that greater AI involvement is more likely to support favorable measurement evidence when HVQ is high.
10.7. Efficiency
AI có theoretical advantage rõ nhất về raw generation.
H7. AI-assisted item-development conditions require less elapsed development time and/or human labor than the human-only condition.
10.8. Efficiency–quality distinction
Thay vì coi đây là một causal hypothesis đơn giản, bài viết xác lập proposition:
P8. Development Efficiency and Measurement Quality constitute theoretically and empirically distinguishable outcome domains.
Điều này cho phép một condition có thể:
high efficiency + high quality;
high efficiency + low quality;
low efficiency + high quality;
low efficiency + low quality.
- PHƯƠNG PHÁP NGHIÊN CỨU ĐỀ XUẤT
11.1. Thiết kế tổng thể
Nghiên cứu đề xuất một:
multi-phase comparative instrument-development experiment.
Quy trình gồm:
- Construct Specification.
- Randomized Item Development.
- Blinded Expert Review.
- Cognitive Interviewing.
- Evidence-Based Revision.
- Pilot Administration.
- Psychometric Evaluation.
- Integrated Validity Argument.
Đơn vị nghiên cứu không nên chỉ là ba final scales.
Để tránh pseudoreplication, mỗi condition nên tạo nhiều independent item pools thông qua independent development teams.
11.2. Construct Specification Dossier
Trước generation, investigators xây dựng dossier gồm:
conceptual definition;
theoretical boundaries;
included facets;
excluded adjacent constructs;
target population;
intended score interpretation;
intended use;
response format;
temporal frame;
language level;
anticipated threats to representation.
Dossier phải giống nhau giữa conditions để tránh confounding.
11.3. Lựa chọn construct
Construct nên:
có theoretical definition đủ rõ;
có multidimensional richness;
có relevance với target population;
không quá phụ thuộc vào một item pool nổi tiếng.
Emerging constructs có thể đặc biệt hữu ích vì giảm nguy cơ LLM tái tạo known scale wording từ training exposure.
11.4. Điều kiện A – Human-Generated
Developers nhận Construct Specification Dossier và standardized item-writing guidance.
Không được sử dụng GenAI.
Ordinary word processor hoặc standardized non-AI references có thể được cho phép theo protocol.
11.5. Điều kiện B – AI-Generated
Cùng dossier được chuyển thành standardized prompt.
AI tạo predetermined number of items.
Trước expert review, human intervention chỉ được phép đối với corrupted output, exact duplicates hoặc formatting errors nếu nghiên cứu cần condition tương đối “pure”.
11.6. Điều kiện C – Human–AI Co-Designed
Human developers tạo facet map hoặc seed framework.
AI được dùng để:
expand coverage;
generate alternatives;
challenge wording;
identify redundancy;
suggest edge cases.
Human developers chịu trách nhiệm candidate pool cuối.
11.7. Optional extension: AI-first versus Human-first
Một extension giàu giá trị lý thuyết là:
AI First → Human Review
so với:
Human First → AI Challenge.
Design này kiểm tra anchoring và quyền khởi tạo conceptual frame.
AI-first có thể neo judgment của human vào machine framing.
Human-first có thể giữ construct architecture trong human control nhưng sử dụng AI như adversarial challenger.
11.8. AI standardization
Methods phải báo đầy đủ:
model name;
provider;
version;
access date;
interface hoặc API;
temperature nếu accessible;
full prompt;
number of generation rounds;
conversation context;
web access;
selection rules;
human edits.
Do LLMs thay đổi theo thời gian, version và access date đặc biệt quan trọng.
- EXPERT REVIEW VÀ CONTENT EVIDENCE
12.1. Panel
Panel có thể gồm khoảng 5–10 experts tùy domain, nhưng không nên xem đây là universal rule.
Panel cần competence complementarity.
12.2. Item-level ratings
Items có thể được đánh giá về:
relevance;
clarity;
construct alignment;
population appropriateness;
potential bias;
redundancy.
12.3. Pool-level ratings
Các tiêu chí pool level gồm:
comprehensiveness;
facet balance;
missing content;
construct contamination.
12.4. Blinding
Experts không biết source condition.
Điều này là thành phần thiết kế trọng yếu để giảm expectancy effects liên quan đến AI.
12.5. Quantitative và qualitative evidence
I-CVI và S-CVI/Ave có thể được sử dụng.
Nhưng expert comments phải được coding và bảo tồn.
Measurement development không nên biến expert review thành một bảng thresholds không có rationale.
- COGNITIVE INTERVIEWING
13.1. Participant selection
Sau expert review, refined items được kiểm tra với members của target population.
Purposive sampling tập trung interpretive variation.
13.2. Interview procedures
Cognitive interviews có thể bao gồm:
think-aloud;
paraphrase;
term interpretation;
retrieval probe;
judgment probe;
response-category explanation;
confidence assessment.
13.3. Coding
Ít nhất hai coders độc lập có thể đánh giá:
correct intended interpretation;
partial interpretation;
systematic misinterpretation;
retrieval difficulty;
construct-irrelevant judgment;
response-mapping difficulty;
context mismatch.
Appropriate inter-rater agreement coefficient được lựa chọn theo measurement level của coded data.
Adjudication diễn ra sau independent coding.
13.4. Revision decisions
Mỗi change được phân loại:
expert-driven;
cognitive-interview-driven;
psychometric-theory-driven;
AI-suggested;
combined evidence.
Version history phải được bảo tồn.
- PILOT TESTING VÀ PSYCHOMETRIC EVALUATION
14.1. Sample-size logic
Pilot sample không nên được xác định duy nhất bằng heuristic “participants per item”.
Sample requirement phụ thuộc:
number of factors;
number of items;
factor loadings;
response categories;
factor correlations;
estimation method;
missingness;
group comparisons;
invariance analyses.
Monte Carlo simulation nên được sử dụng khi CFA/SEM design đủ phức tạp.
14.2. Descriptive item analysis
Báo cáo:
missingness;
category frequencies;
floor/ceiling patterns;
item distributions;
response time khi có ý nghĩa.
14.3. Dimensionality
Có thể sử dụng:
parallel analysis;
EFA;
CFA;
ESEM khi substantively justified.
Không dùng model fit thresholds như absolute laws.
14.4. Reliability/precision
Báo omega hoặc model-based reliability khi appropriate.
Alpha có thể báo thêm cho comparability.
14.5. Relations to external variables
External variables nên được prespecify từ theory.
Không chọn correlations thuận lợi sau khi đã xem data.
14.6. IRT và item functioning
Nếu design cho phép:
item difficulty;
discrimination;
information functions;
DIF
có thể cung cấp additional evidence.
- STATISTICAL ANALYSIS PLAN
15.1. Confirmatory focus
Primary confirmatory focus nên được prespecify trước data collection.
Đối với study này, một cách cấu hình hợp lý là:
Primary experimental contrast:
AI-Generated vs. Human–AI Co-Designed.
Contrast này giữ AI involvement tương đối cao ở cả hai điều kiện nhưng thay đổi mức độ human co-design/verification.
Primary upstream outcome:
Content Validity Quality.
Key downstream outcomes:
Response-Process Quality và prespecified components của PMQ.
DE được phân tích như parallel outcome.
15.2. H1
Expert ratings thường nested trong items và items nested trong pools/teams.
Multilevel models vì vậy phù hợp hơn aggregation đơn giản.
Model có thể gồm random intercepts cho item/pool và expert tùy crossing structure.
15.3. H2–H4
CVQ, RPQ và PMQ phải được phân tích ở level tương thích.
Không nên aggregate item-level evidence lên một final scale quá sớm.
15.4. H5 – mediation
Mediation chỉ được estimate nếu:
có nhiều independently developed pools;
CVQ được xác định trước RPQ;
RPQ được đo trước downstream psychometric evaluation;
sample size ở pool/item level đủ để estimate indirect effect.
Bootstrap hoặc simulation-based uncertainty có thể được sử dụng tùy model.
15.5. H6 – moderation
HVQ moderation nên được kiểm tra trong AI-involved conditions.
Một design mạnh hơn là experimentally manipulate verification architecture thay vì chỉ quan sát natural variation.
15.6. H7 và P8
DE outcomes gồm:
elapsed time;
human person-hours;
generation rounds;
revision rounds;
expert-ready yield;
usable-items-per-hour;
API/financial cost.
P8 có thể được kiểm tra bằng cách xem efficiency và measurement-quality metrics có tạo patterns phân biệt giữa conditions hay không.
Một efficiency–quality frontier có thể hữu ích hơn một correlation duy nhất.
15.7. Multiplicity
Confirmatory hypotheses phải được phân biệt exploratory analyses.
Nếu nhiều PMQ outcomes được kiểm định đồng thời, multiplicity strategy phải được prespecify.
Không nên tạo một chuỗi post hoc tests và chỉ báo outcomes thuận lợi.
15.8. Missing data
Missingness được mô tả theo source và stage.
Handling strategy phải tương thích với assumptions về missing-data mechanism.
Complete-case analysis không nên là mặc định.
15.9. Sensitivity analyses
Có thể gồm:
alternative factor structures;
alternative item-retention rules;
robust estimators;
inclusion/exclusion of borderline items;
alternative operationalization của CVQ/RPQ;
alternative coding thresholds;
models có hoặc không điều chỉnh developer expertise.
15.10. Preregistration
Primary hypotheses, experimental contrasts, item-retention principles và Statistical Analysis Plan nên được preregister khi khả thi.
AI prompts và decision logs cũng nên được archived nếu privacy/IP cho phép.
Reproducibility đòi hỏi nhiều hơn câu:
“We used ChatGPT.”
- DEVELOPMENT EFFICIENCY
Efficiency cần được định nghĩa rõ.
Elapsed time và human labor không đồng nhất.
Một LLM có thể tạo pool trong vài giây nhưng đòi hỏi hàng giờ verification.
Do đó DE nên được phân tích theo nhiều resource dimensions.
Một useful-item-per-hour metric có thể cung cấp thông tin hơn raw generation speed.
Mục tiêu không phải chứng minh AI nhanh.
Điều đó khá hiển nhiên ở generation stage.
Câu hỏi có giá trị hơn là:
AI có giảm tổng resource cost sau khi tính cả verification và rework hay không?
- PLANNED RESULTS AND ANALYTICAL REPORTING FRAMEWORK
Bản thảo hiện tại là theoretical-development / experimental-protocol article.
Do chưa tiến hành empirical study, bài viết không báo cáo:
β;
p;
confidence intervals;
CFI/TLI;
RMSEA;
factor loadings;
reliability coefficients;
AVE;
HTMT;
effect sizes giả.
Khi nghiên cứu được thực hiện, Results nên theo trình tự:
17.1. Sample và process fidelity
Báo participant characteristics, randomization, adherence và deviations.
Kiểm tra Human condition có thực sự không dùng GenAI.
Kiểm tra AI conditions có tuân protocol.
17.2. Efficiency outcomes
Báo:
elapsed time;
person-hours;
generation rounds;
revision rounds;
usable-item yield.
17.3. Content-validity evidence
Báo:
I-CVI distribution;
pool-level coverage;
expert concerns;
condition differences.
17.4. Cognitive-interview findings
Báo:
comprehension failures;
retrieval problems;
judgment mismatch;
response-mapping problems;
revision rate.
Không chỉ báo “90% items clear”.
Cần mô tả loại lỗi.
17.5. Item trajectories
Một contribution đặc biệt hữu ích là báo item lineage:
AI-generated → rejected
AI-generated → revised → retained
AI-generated → retained unchanged
Human-generated → rejected
Human–AI → cognitively revised → retained.
Trajectories cho biết source tương tác với evidence như thế nào.
17.6. Psychometric evidence
Báo dimensionality, loadings, uncertainty, fit, residual problems, reliability và external relations.
17.7. Mediation/moderation
Báo point estimates, confidence intervals và effect sizes.
Không diễn giải chỉ dựa vào p < .05.
- THẢO LUẬN
18.1. Framework giải thích điều gì?
Framework giải thích một nghịch lý mới của measurement science:
AI có thể đồng thời làm item production tốt hơn và khiến validation shortcuts dễ xảy ra hơn.
Generative capability làm giảm scarcity của candidate items nhưng không làm giảm scarcity của high-quality evidence.
Do đó, novelty của lĩnh vực phải chuyển từ câu hỏi về generation sang câu hỏi về evidential transition.
AI Item Generation ≠ Instrument Validation.
18.2. Vì sao distinction này quan trọng về lý thuyết?
Measurement theory vốn đã phân biệt content, response processes, internal structure và external relations.
GenAI không thay đổi logic validity cơ bản.
Nó thay đổi risk architecture.
Khi production từng rất tốn kém, researcher có động lực đầu tư cẩn thận vào từng item.
Khi AI có thể tạo hàng trăm items tức thời, temptation chuyển thành:
generate broadly;
screen rapidly;
run CFA;
retain best-fitting items.
Điều này có thể tạo một form mới của psychometric optimization without construct optimization.
Framework vì vậy đặt evidence governance vào trung tâm.
18.3. Linguistic plausibility
AI-generated text có thể tạo professional appearance mạnh.
Trong scientific measurement, chính sự trôi chảy này có thể là risk.
Researcher có thể underestimate:
construct drift;
cultural assumptions;
semantic redundancy;
hidden ambiguity.
Do đó:
Linguistic Quality ≠ Measurement Quality.
18.4. Expert evidence và respondent evidence
Expert evaluation và cognitive interviewing phải được nối nhưng không hòa lẫn.
Expert hỏi:
“Should the item mean X?”
Respondent evidence hỏi:
“Does the item actually evoke X?”
Một scale mạnh cần cả theoretical representation lẫn actual response-process alignment.
18.5. Internal structure không phải endpoint
Một bài scale-development không nên kết thúc validity discussion sau CFA.
Internal structure có thể rất mạnh nhưng không chữa được content flaw.
Đây là lý do PMQ được xem như downstream evidence profile, không phải proxy cho toàn bộ validity.
18.6. Human Verification Quality như cơ chế
Một đóng góp lý thuyết quan trọng của bài là chuyển human involvement từ binary presence/absence thành verification quality.
Câu hỏi không phải:
“Có người tham gia không?”
Mà là:
Human kiểm tra gì?
Dựa trên evidence nào?
Output nào bị reject?
Revision được biện minh thế nào?
Uncertainty có được ghi lại không?
Decision chain có audit được không?
HVQ biến human agency từ một normative slogan thành methodological mechanism.
18.7. Human–AI co-design không mặc nhiên tốt nhất
Framework không dự đoán Human–AI luôn superior.
Kết quả phụ thuộc configuration of roles.
AI phù hợp với:
breadth generation;
linguistic alternatives;
redundancy suggestions;
adversarial critique.
Humans phù hợp với:
construct boundaries;
contextual meaning;
ethical judgment;
evidence synthesis;
final interpretation;
accountability.
Nếu humans rubber-stamp machine outputs, “co-design” về tên gọi nhưng thực chất gần automation dependence.
Nếu AI chỉ sửa chính tả, workflow lại không khai thác generative capability.
Do đó, explanatory variable thực sự có thể là architecture of collaboration, không đơn thuần access to AI.
18.8. Efficiency–quality frontier
AI-assisted development cần được đánh giá trên ít nhất hai dimensions:
Efficiency
và
Measurement Quality.
Điều này tạo bốn configurations:
High efficiency / High quality: augmentation.
High efficiency / Low quality: acceleration only.
Low efficiency / High quality: quality-oriented but resource intensive.
Low efficiency / Low quality: unfavorable configuration.
Framework này giúp tránh productivity fallacy.
18.9. Generative psychometrics và evidence governance
AI-GENIE đại diện bước phát triển quan trọng của generative psychometrics khi nối generation với network psychometric evaluation.
Framework hiện tại bổ sung một layer khác:
Generation
→ Content Warrant
→ Response-Process Warrant
→ Psychometric Warrant
→ Integrated Validity Argument
→ Human Accountability.
Generative psychometrics do đó không nên chỉ phát triển algorithms tạo items tốt hơn.
Nó cần phát triển evidence governance.
- HÀM Ý PHƯƠNG PHÁP
Một workflow cần tránh là:
Generate → CFA → Validate.
Quy trình tối thiểu được đề xuất:
Construct Definition
→ Domain Mapping
→ Human/AI Item Generation
→ Human Construct Verification
→ Blinded Expert Review
→ Content Evidence
→ Cognitive Interviewing
→ Response-Process Evidence
→ Evidence-Based Revision
→ Pilot Testing
→ Internal-Structure Evidence
→ Reliability/Precision
→ Relations to Other Variables
→ Integrated Validity Argument.
AI nên tạo rộng.
Humans nên xác định construct boundaries.
Một prompt không nên yêu cầu:
“Create a validated scale.”
Nên yêu cầu:
“Generate candidate items for each predefined facet. Identify potential ambiguity, redundancy, and overlap with adjacent constructs for human evaluation. Do not claim that the items are valid.”
Cách prompt này giới hạn epistemic overclaim ngay từ đầu.
- ITEM PROVENANCE VÀ REPRODUCIBILITY
Future AI-assisted scale-development papers nên báo:
source condition;
original AI output;
model/version;
prompt;
human edit;
reason for edit;
expert evidence;
cognitive evidence;
pilot evidence;
final wording.
Item provenance giúp reviewer biết một final item được hình thành qua những quyết định nào.
Điều này đặc biệt quan trọng vì cùng một model có thể thay đổi theo thời gian.
Reproducibility trong AI research không còn chỉ là chia sẻ final questionnaire.
Nó cần chia sẻ process architecture khi ethics và IP cho phép.
- HÀM Ý CHO ĐÀO TẠO PHƯƠNG PHÁP NGHIÊN CỨU
GenAI không làm psychometric competence ít cần thiết.
Nó làm competence quan trọng hơn.
Khi production barrier giảm, validation responsibility vẫn giữ nguyên.
Graduate students không chỉ cần học:
item writing;
CVI;
EFA;
CFA;
reliability.
Họ còn cần học:
construct drift;
prompt-to-item risk;
synthetic redundancy;
response-process verification;
AI disclosure;
provenance;
data privacy;
human accountability.
Một nhà nghiên cứu có khả năng dùng AI nhưng không có measurement competence có thể tạo questionnaire trông rất chuyên nghiệp nhưng không có evidential warrant.
- ĐẠO ĐỨC, TRANSPARENCY VÀ AI GOVERNANCE
22.1. AI như research method
Nếu GenAI trực tiếp tạo candidate research items, AI không chỉ là writing assistant.
Nó là một phần của research method.
Vì vậy, model, version, prompt, selection process và human oversight phải được báo trong Methods.
22.2. Privacy
Nếu transcripts, participant responses hoặc identifiable research data được đưa vào external AI services, cần xem xét:
informed consent;
data minimization;
provider retention policy;
institutional rules;
research ethics approval;
intellectual property.
Không nên mặc định commercial AI system tương đương secure research environment.
22.3. Bias và fairness
AI-generated items có thể phản ánh linguistic hoặc cultural bias từ training data.
Fairness analysis nên cân nhắc:
representation;
stereotypes;
accessibility;
DIF;
cross-group interpretation.
Machine generation không tạo automated neutrality.
22.4. Authorship và accountability
AI không chịu trách nhiệm về:
construct definition;
measurement decisions;
validity claims;
ethical compliance.
Human investigators phải giữ final accountability.
Do đó:
Automation of Generation ≠ Delegation of Accountability.
- HẠN CHẾ
23.1. Model dependence
Kết quả từ một model không generalize sang mọi LLM.
Performance phụ thuộc:
model architecture;
training data;
alignment;
version;
prompt;
temperature;
language.
23.2. Construct dependence
AI có thể hoạt động đặc biệt tốt với well-established constructs vì training corpus chứa nhiều descriptions.
Emerging hoặc culturally bounded constructs có thể khó hơn.
23.3. Expertise dependence
Novice và expert sử dụng cùng AI có thể tạo results rất khác.
Future studies nên so sánh:
Novice Human;
Expert Human;
Novice + AI;
Expert + AI.
23.4. Language dependence
English-language performance không bảo đảm Vietnamese-language quality.
Future designs nên so sánh:
Vietnamese-first generation;
English-first + translation;
bilingual co-design;
back-translation;
cognitive equivalence;
measurement invariance.
23.5. AI-first anchoring
Nếu AI khởi tạo conceptual frame, human developers có thể bị anchor.
Human-First → AI Challenge versus AI-First → Human Review là hướng thực nghiệm quan trọng.
23.6. Short-term validation
Initial CFA success không bảo đảm temporal stability, predictive performance hoặc invariance over time.
AI-generated measures cần longitudinal evidence giống human-developed instruments.
- HƯỚNG NGHIÊN CỨU TƯƠNG LAI
Generative psychometrics có thể mở rộng sang:
automated item banks;
IRT-calibrated generation;
adaptive testing;
dynamic assessment;
multilingual item generation;
AI-assisted DIF detection.
Nhưng mức tự động hóa càng cao, evidence governance càng quan trọng.
Future research nên kiểm tra không chỉ:
Can AI generate better items?
mà còn:
Can AI-assisted systems generate better validity arguments?
Một research program dài hạn có thể theo chuỗi:
Generative Capability
→ Verification Architecture
→ Evidence Quality
→ Measurement Decisions
→ Scientific Accountability.
- SÁU MỆNH ĐỀ HỌC THUẬT TRUNG TÂM
P1. AI Item Generation ≠ Instrument Validation.
Generative capability là upstream production capacity.
P2. Linguistic Quality ≠ Measurement Quality.
Fluency không chứng minh construct representation.
P3. Expert Agreement ≠ Response-Process Validity Evidence.
Expert intention không phải respondent cognition.
P4. Good Model Fit ≠ Valid Measurement.
Internal structure là một nguồn validity evidence.
P5. Human Review ≠ Methodological Verification.
Human presence chỉ tạo methodological value khi tạo traceable evidential checks.
P6. Efficiency Gain ≠ Measurement Gain.
Development productivity và epistemic quality là outcome domains riêng biệt.
Sáu mệnh đề tạo thành chuỗi:
Generation Capability
→ Representation Risk
→ Multi-Source Evidence
→ Human Verification
→ Psychometric Evaluation
→ Integrated Validity Argument
→ Scientific Accountability.
- ĐÓNG GÓP LÝ THUYẾT
26.1. Generation–Validation Distinction
Bài viết chuyển debate khỏi binary Human versus AI.
Câu hỏi quan trọng không phải ai viết item.
Câu hỏi là evidence nào cho phép sử dụng responses để suy luận.
26.2. Multi-Source Validity Architecture
CVQ, RPQ và PMQ được tách thành ba evidential layers liên quan nhưng không đồng nhất.
26.3. Human Verification Theory
HVQ biến generic human-in-the-loop rhetoric thành construct có khả năng operationalize.
26.4. Efficiency–Quality Separation
Development productivity được tách khỏi scientific measurement improvement.
26.5. Traceable Item Provenance
Mỗi item được xem là một evidential object có lịch sử phát triển có thể audit.
26.6. Evidence-Governed Generative Psychometrics
Đóng góp tổng hợp của bài là đề xuất:
Evidence-Governed Generative Psychometrics
trong đó generative capability phải được nối với content warrant, respondent cognition, psychometric evidence, verification và human accountability.
- KẾT LUẬN
GenAI đã thay đổi một ràng buộc lâu đời của measurement science.
Trước đây, xây dựng một large candidate-item pool đòi hỏi đáng kể human labor. Hiện nay, một LLM có thể tạo hàng chục hoặc hàng trăm candidate items gần như tức thời.
Evidence mới cho thấy lợi ích này không chỉ là tốc độ. AI-generated items trong một số contexts có thể đạt internal structure và reliability cạnh tranh với traditionally developed measures.
Nhưng khả năng đó không làm validation trở nên ít cần thiết.
Nó làm validation quan trọng hơn.
Khi generation trở nên rẻ, dễ và gần như không giới hạn, bottleneck của measurement science dịch chuyển từ:
PRODUCTION
sang:
VERIFICATION.
Một item không trở nên scientifically warranted vì AI tạo nó.
Một item cũng không trở nên scientifically warranted chỉ vì một human expert viết nó.
Source ≠ Validity Evidence.
Điều quan trọng là:
item có đại diện construct không;
pool có bao phủ domain không;
item có tránh construct contamination không;
target respondents có hiểu item như intended không;
response processes có phù hợp intended interpretation không;
empirical structure có tương thích theory không;
scores có relations với external variables như dự kiến không;
và research team có giải trình được toàn bộ decision chain không.
Bài viết vì vậy đề xuất architecture:
ITEM DEVELOPMENT CONDITION
→ CONTENT VALIDITY QUALITY
→ RESPONSE-PROCESS QUALITY
→ PSYCHOMETRIC MEASUREMENT QUALITY
→ DEFENSIBLE VALIDITY ARGUMENT
với:
HUMAN VERIFICATION QUALITY
là boundary condition quyết định AI assistance trở thành:
AUGMENTATION
hay chỉ:
ACCELERATION,
và:
DEVELOPMENT EFFICIENCY
là một outcome domain riêng biệt.
Luận đề cuối cùng của bài là:
Giá trị khoa học của việc phát triển bảng hỏi có hỗ trợ AI không nên được đánh giá dựa trên số lượng mục hỏi mà một mô hình AI tạo sinh có thể tạo ra, mức độ trôi chảy về ngôn ngữ của các mục hỏi đó, hay tốc độ hoàn thiện một thang đo. Thay vào đó, giá trị này cần được đánh giá dựa trên việc quá trình phát triển có hỗ trợ AI có đóng góp vào việc hình thành một chuỗi bằng chứng về giá trị đo lường minh bạch, có khả năng truy nguyên, được con người xác minh và có thể bảo vệ bằng bằng chứng thực nghiệm hay không; qua đó hỗ trợ một cách có cơ sở cho các diễn giải và cách sử dụng điểm số đã được xác định.
Nói cách khác:
AI có thể tự động hóa việc tạo mục hỏi; nhưng không thể tự động hóa trách nhiệm khoa học đối với ý nghĩa của điểm số.
TUYÊN BỐ VỀ VIỆC SỬ DỤNG AI TẠO SINH VÀ CÁC CÔNG NGHỆ HỖ TRỢ AI TRONG QUÁ TRÌNH CHUẨN BỊ BẢN THẢO
Trong quá trình chuẩn bị bản thảo này, tác giả/nhóm tác giả đã sử dụng [TÊN CÔNG CỤ/DỊCH VỤ AI] để hỗ trợ [hiệu chỉnh ngôn ngữ / tổ chức cấu trúc bản thảo / tổng hợp tài liệu / mục đích cụ thể khác]. Sau khi sử dụng các công cụ này, tác giả/nhóm tác giả đã tiến hành phản biện, kiểm chứng độc lập và hiệu chỉnh các nội dung có sự hỗ trợ của AI khi cần thiết. Tác giả/nhóm tác giả chịu hoàn toàn trách nhiệm về tính chính xác, tính toàn vẹn, các diễn giải, trích dẫn và nội dung cuối cùng của bản thảo.
Trong trường hợp AI tạo sinh hoặc các công nghệ hỗ trợ AI được sử dụng như một bộ phận cấu thành của phương pháp nghiên cứu thực nghiệm — chẳng hạn để tạo các mục hỏi ứng viên cho bảng hỏi hoặc thang đo — việc sử dụng AI cần được báo cáo một cách minh bạch và có khả năng tái lập trong phần Phương pháp nghiên cứu (Methods). Tùy theo đặc điểm của nghiên cứu, thông tin báo cáo cần bao gồm mô hình AI và nhà cung cấp, phiên bản mô hình, ngày truy cập, câu lệnh hoặc giao thức tạo câu lệnh (prompting protocol), kiến trúc tương tác, các thiết lập hệ thống có liên quan, quy trình tạo và lựa chọn đầu ra, các hiệu chỉnh do con người thực hiện, cũng như các thủ tục xác minh của con người (human-verification procedures). Các đầu ra do AI tạo không được xem là bằng chứng nghiên cứu đã được xác nhận độc lập nếu chưa trải qua các quy trình kiểm chứng phương pháp và xác minh của con người phù hợp.
Tác giả/nhóm tác giả chịu hoàn toàn trách nhiệm đối với mọi phán đoán khoa học, quyết định phương pháp luận, việc kiểm chứng nguồn và trích dẫn, các diễn giải và kết luận được trình bày trong bản thảo.
TÀI LIỆU THAM KHẢO
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
- American Psychological Association. (2020). APA guidelines for psychological assessment and evaluation. American Psychological Association.
- Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R., & Young, S. L. (2018). Best practices for developing and validating scales for health, social, and behavioral research: A primer. Frontiers in Public Health, 6, 149. https://doi.org/10.3389/fpubh.2018.00149
- Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061–1071. https://doi.org/10.1037/0033-295X.111.4.1061
- Clark, L. A., & Watson, D. (2019). Constructing validity: New developments in creating objective measuring instruments. Psychological Assessment, 31(12), 1412–1427. https://doi.org/10.1037/pas0000626
- DeVellis, R. F., & Thorpe, C. T. (2022). Scale development: Theory and applications (5th ed.). SAGE.
- Downing, S. M. (2003). Validity: On the meaningful interpretation of assessment data. Medical Education, 37(9), 830–837. https://doi.org/10.1046/j.1365-2923.2003.01594.x
- (2026). Generative AI policies for journals. Elsevier.
- González Canché, M. S. (2026). Data science, interactive visualizations, and generative AI tools for the analysis of qualitative, mixed-methods, and multimodal evidence. Morgan Kaufmann.
- Haynes, S. N., Richard, D. C. S., & Kubany, E. S. (1995). Content validity in psychological assessment: A functional approach to concepts and methods. Psychological Assessment, 7(3), 238–247. https://doi.org/10.1037/1040-3590.7.3.238
- Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43, 115–135. https://doi.org/10.1007/s11747-014-0403-8
- Hoffmann, S., Lasarov, W., & Dwivedi, Y. K. (2024). AI-empowered scale development: Testing the potential of ChatGPT. Technological Forecasting and Social Change, 205, 123488. https://doi.org/10.1016/j.techfore.2024.123488
- Kuru, H. (2025). Rethinking survey development in health research with AI-driven methodologies. Frontiers in Digital Health, 7, 1636333. https://doi.org/10.3389/fdgth.2025.1636333
- Lazar, J., Feng, J. H., & Hochheiser, H. (2017). Research methods in human-computer interaction (2nd ed.). Morgan Kaufmann.
- MacDermid, J. C. (2021). ICF linking and cognitive interviewing are complementary methods for optimizing content validity of outcome measures: An integrated methods review. Frontiers in Rehabilitation Sciences, 2, 702596. https://doi.org/10.3389/fresc.2021.702596
- McNeish, D. (2018). Thanks coefficient alpha, we’ll take it from here. Psychological Methods, 23(3), 412–433. https://doi.org/10.1037/met0000144
- Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
- Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO.
- Mokkink, L. B., Elsman, E. B. M., & Terwee, C. B. (2024). COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Quality of Life Research, 33, 2929–2939. https://doi.org/10.1007/s11136-024-03761-6
- Nisbet, R., Miner, G. D., & McCormick, K. (2024). Handbook of statistical analysis: AI and ML applications (3rd ed.). Academic Press.
- Polit, D. F., & Beck, C. T. (2006). The content validity index: Are you sure you know what’s being reported? Critique and recommendations. Research in Nursing & Health, 29(5), 489–497. https://doi.org/10.1002/nur.20147
- Russell-Lasalandra, L. L., Christensen, A. P., & Golino, H. (2026). Generative psychometrics via AI-GENIE: Automatic item generation and validation with network-integrated evaluation. Behavior Research Methods, 58, Article 217. https://doi.org/10.3758/s13428-026-03082-1
- Salah, M., Abdelfattah, F., Al Halbusi, H., Jassem, S., Mohammed, M., Ismail, M. M., & Al Balghouni, A. (2025). Can generative AI craft scale items? A mixed-method study on AI’s capability to adapt and create new scales with recommendations for best practices. Social Sciences & Humanities Open, 12, 101698. https://doi.org/10.1016/j.ssaho.2025.101698
- Terwee, C. B., Prinsen, C. A. C., Chiarotto, A., Westerman, M. J., Patrick, D. L., Alonso, J., Bouter, L. M., de Vet, H. C. W., & Mokkink, L. B. (2018). COSMIN methodology for evaluating the content validity of patient-reported outcome measures: A Delphi study. Quality of Life Research, 27, 1159–1170. https://doi.org/10.1007/s11136-018-1829-0
- Terry, J., Strait, G., Alsarraf, S., Weinmann, E., et al. (2025). Artificial intelligence in scale development: Evaluating AI-generated survey items against gold standard measures. Current Psychology, 44, 16339–16350. https://doi.org/10.1007/s12144-025-08240-w
- Valenzuela, S., Winter, S., & Rivera, S. (2025). Using large language models for survey research in communication: Opportunities and challenges. Communication and Change, 1, Article 14. https://doi.org/10.1007/s44382-025-00014-z
- Williamson, K., & Johanson, G. (Eds.). (2018). Research methods: Information, systems, and contexts (2nd ed.). Chandos Publishing.
- Willis, G. B. (2005). Cognitive interviewing: A tool for improving questionnaire design. SAGE.
