From Methodological Knowledge to Valid Scientific Inference: The Roles of Methodological Fit, Research Design Quality, and AI Literacy in Quantitative Research
NGUYỄN VĂN HÙNG
HIGHLIGHTS
- Phát triển chuỗi lý thuyết Methodological Knowledge → Methodological Fit → Research Design Quality → Scientific Inference Quality.
- Khái niệm hóa Methodological Fit như một relational methodological competence thay vì chỉ là kiến thức về kỹ thuật thống kê.
- Phân biệt rõ Methodological Fit với Research Design Quality, qua đó tránh construct contamination và tautological relationships.
- Định vị Research Design Quality như cơ chế chuyển methodological judgment thành evidence-generating architecture.
- Khái niệm hóa Scientific Inference Quality như một epistemic outcome tách biệt với statistical significance và analytical sophistication.
- Định vị AI Literacy như boundary condition/capability amplifier của quan hệ MF → RDQ, thay vì mặc định là predictor trực tiếp của research quality.
- Đề xuất Human Methodological Accountability như nguyên tắc quản trị cốt lõi của nghiên cứu định lượng có AI hỗ trợ.
TÓM TẮT
Bối cảnh
Generative Artificial Intelligence (GenAI) đang làm thay đổi sâu sắc cách người học và nhà nghiên cứu thực hiện nghiên cứu định lượng. Các hệ thống AI hiện có thể hỗ trợ xác lập giả thuyết, đề xuất thiết kế, tính cỡ mẫu, lựa chọn kỹ thuật thống kê, tạo mã R/Python, kiểm tra giả định, trực quan hóa dữ liệu và diễn giải kết quả. Việc giảm mạnh rào cản kỹ thuật tạo ra cơ hội dân chủ hóa năng lực phân tích, nhưng đồng thời đặt ra một vấn đề nhận thức luận quan trọng: khả năng thực hiện phân tích không đồng nghĩa với khả năng thiết kế nghiên cứu và đưa ra suy luận khoa học có giá trị.
Vấn đề nghiên cứu
Năng lực nghiên cứu định lượng thường được tiếp cận thông qua kiến thức thống kê hoặc khả năng thực hiện kỹ thuật. Cách tiếp cận này chưa phân biệt đầy đủ giữa việc biết một phương pháp tồn tại và khả năng phán đoán phương pháp nào phù hợp với một vấn đề khoa học cụ thể. Trong môi trường AI, khoảng cách này càng quan trọng bởi AI có thể tạo ra những đề xuất kỹ thuật có vẻ hợp lý về mặt ngôn ngữ nhưng không nhất thiết phù hợp với câu hỏi nghiên cứu, construct, measurement, sampling mechanism, design, estimand hoặc phạm vi suy luận.
Mục tiêu
Bài viết phát triển mô hình tích hợp:
Methodological Knowledge (MK) → Methodological Fit (MF) → Research Design Quality (RDQ) → Scientific Inference Quality (SIQ)
trong đó AI Literacy (AIL) được lý thuyết hóa như một boundary condition của quan hệ MF → RDQ.
MK đại diện cho methodological repertoire; MF đại diện cho context-sensitive methodological judgment; RDQ phản ánh chất lượng của evidence-generating architecture; và SIQ phản ánh mức độ scientific claims được hiệu chỉnh tương xứng với evidence.
Phương pháp đề xuất
Nghiên cứu sử dụng multi-method observational design kết hợp structured measures và performance assessment. MK được đánh giá thông qua knowledge-, scenario-, error-detection- và inference-boundary tasks. MF, RDQ và SIQ được đánh giá bằng performance artifacts cùng behaviorally anchored rubrics bởi những người chấm độc lập và làm mù. AIL được đo bằng thang đo có bằng chứng giá trị phù hợp, bổ sung bằng performance indicators về AI recommendation verification, source verification, assumption checking và code validation. Structural equation modeling và conditional process analysis được đề xuất để kiểm định H1–H7.
Đóng góp
Bài viết chuyển trọng tâm từ statistical competence sang methodological alignment competence, phân biệt knowledge, judgment, design architecture và epistemic outcome. Ba mệnh đề trung tâm được đề xuất:
Sophisticated Analysis ≠ Good Research Design.
AI Recommendation ≠ Methodological Justification.
Statistical Significance ≠ Valid Scientific Inference.
AI Literacy vì vậy không được xem như nguồn methodological competence độc lập, mà như năng lực có thể khuếch đại quá trình chuyển methodological judgment thành research design có chất lượng, với điều kiện con người duy trì trách nhiệm kiểm chứng và biện minh.
Từ khóa: methodological knowledge; methodological fit; research design quality; scientific inference quality; AI literacy; generative AI; quantitative research; methodological judgment.

ABSTRACT
Background
Generative artificial intelligence is rapidly lowering the technical barriers associated with quantitative research. Contemporary AI systems can recommend analytical procedures, generate executable statistical code, support sample-size planning, check assumptions, visualize data, and draft interpretations of statistical results. These capabilities can increase analytical accessibility and efficiency. However, technical accessibility does not necessarily imply methodological appropriateness, and analytical sophistication does not guarantee valid scientific inference.
Research problem
Quantitative research competence is frequently conceptualized in terms of statistical knowledge or technical execution. Such conceptualizations insufficiently distinguish between knowing available methods and judging which methodological configuration is defensible for a particular scientific problem. This distinction becomes increasingly consequential when GenAI can generate technically plausible methodological recommendations without independently establishing their epistemic appropriateness.
Purpose
This article develops an integrated theoretical model linking Methodological Knowledge (MK), Methodological Fit (MF), Research Design Quality (RDQ), and Scientific Inference Quality (SIQ). AI Literacy (AIL) is theorized as a boundary condition affecting the translation of methodological fit into research design quality.
Proposed method
A multi-method observational design combining structured measurement and performance assessment is proposed. MK is assessed through knowledge- and scenario-based tasks. MF, RDQ, and SIQ are assessed through research artifacts using behaviorally anchored rubrics scored by independent blinded raters. AIL is measured using an instrument supported by appropriate validity evidence and supplemented with performance indicators of AI evaluation and verification. Structural equation modeling and conditional process analysis are proposed to evaluate direct, indirect, moderating, and conditional indirect relationships.
Contribution
The framework shifts the conceptual focus from statistical competence toward methodological alignment competence. It distinguishes methodological repertoire from methodological judgment, research architecture, and epistemic outcomes while positioning AI Literacy as a conditional capability amplifier rather than an inherently beneficial predictor of research quality. Three propositions summarize the theoretical contribution: Sophisticated Analysis ≠ Good Research Design; AI Recommendation ≠ Methodological Justification; Statistical Significance ≠ Valid Scientific Inference.
Keywords: methodological knowledge; methodological fit; research design quality; scientific inference quality; AI literacy; quantitative research; generative AI; methodological judgment.
- GIỚI THIỆU
1.1. GenAI và sự thay đổi cấu trúc của năng lực nghiên cứu định lượng
GenAI đang làm thay đổi đáng kể quá trình sản xuất tri thức khoa học. Trong nghiên cứu định lượng, nhiều hoạt động từng đòi hỏi kiến thức phần mềm, kinh nghiệm lập trình hoặc khả năng thao tác thống kê chuyên biệt ngày càng có thể được thực hiện thông qua giao diện ngôn ngữ tự nhiên. Người dùng có thể yêu cầu AI đề xuất statistical test, tạo code, giải thích assumptions, tính effect size, xây visualization hoặc diễn giải model output.
Sự thay đổi này có thể làm tăng đáng kể analytical accessibility. Tuy nhiên, nó đồng thời làm suy yếu giá trị của technical execution như một proxy cho research competence. Khi code có thể được tạo trong vài giây, câu hỏi quan trọng không còn chỉ là researcher có thể chạy model hay không, mà là researcher có hiểu vì sao model đó phù hợp, assumptions nào phải được bảo vệ, design nào tạo ra dữ liệu và claim nào evidence thực sự hỗ trợ hay không.
Do đó, sự phổ biến của GenAI không làm methodological knowledge và judgment trở nên ít quan trọng. Ngược lại, khi execution được tự động hóa, selection, evaluation, verification và justification trở thành những năng lực có giá trị phân biệt lớn hơn.
Luận điểm đầu tiên của bài viết là:
AI có thể làm cho sophisticated analysis dễ thực hiện hơn, nhưng không làm methodological fit trở thành tự động.
1.2. Technical correctness và methodological appropriateness
Một phân tích có thể technically correct nhưng methodologically inappropriate.
Một regression model có thể chạy đúng cú pháp nhưng không trả lời causal question.
Một SEM có thể tạo satisfactory fit indices nhưng dựa trên construct representation yếu.
Một sample rất lớn có thể tạo narrow confidence intervals nhưng vẫn chứa selection bias nghiêm trọng.
Một statistically significant coefficient có thể không có substantive importance.
Một AI-generated explanation có thể trôi chảy về ngôn ngữ nhưng vượt quá evidence.
Những trường hợp này cho thấy cần phân biệt:
Correct Analysis
với
Appropriate Analysis.
Correct analysis liên quan tới việc procedure được thực hiện chính xác.
Appropriate analysis liên quan tới việc procedure có phù hợp với research question, design, data-generating process, estimand, assumptions và intended inference hay không.
Vì vậy:
Sophisticated Analysis ≠ Good Research Design.
1.3. Khoảng cách giữa knowledge và judgment
Methodological education thường được tổ chức theo repertoire các kỹ thuật: descriptive statistics, t-test, ANOVA, regression, factor analysis, SEM và các kỹ thuật mở rộng.
Cấu trúc này cần thiết nhưng chưa đủ.
Một researcher có thể biết regression assumptions nhưng vẫn lựa chọn regression không phù hợp với research question.
Người học có thể biết định nghĩa representative sampling nhưng vẫn khái quát hóa từ convenience sample sang một population quá rộng.
Họ có thể hiểu reliability nhưng sử dụng một measure không đại diện đầy đủ cho construct.
Do đó, methodological competence phải bao gồm ít nhất hai tầng:
Methodological Knowledge: biết principles, methods, assumptions và alternatives.
Methodological Fit: biết phương pháp nào phù hợp với scientific problem và có thể biện minh lựa chọn đó.
Nếu MK trả lời:
What methodological options exist?
MF trả lời:
Which option is justified here, and why?
1.4. Research design như evidence-generating architecture
Scientific evidence không được tạo bởi statistical analysis một cách độc lập.
Trước khi dataset được phân tích, researcher đã đưa ra một chuỗi quyết định:
Question → Construct → Measurement → Population → Sampling → Design → Data → Analysis.
Những quyết định upstream đặt giới hạn cho evidence downstream.
Nếu measurement không đại diện construct, statistical sophistication không tạo content validity.
Nếu sampling mechanism không hỗ trợ generalization, N lớn không tự mở rộng target population.
Nếu design không hỗ trợ causal identification, regression không tự biến association thành causality.
Do đó, research design phải được hiểu như một evidence-generating architecture.
1.5. Scientific inference như epistemic outcome
Mục tiêu cuối cùng của quantitative research không phải tạo p-value, coefficient hay fit index.
Mục tiêu là tạo ra scientific claims có mức độ mạnh tương xứng với evidence.
Scientific Inference Quality vì vậy liên quan đến:
evidence–claim proportionality;
uncertainty recognition;
causal restraint;
generalization appropriateness;
alternative explanations;
boundary conditions.
Nguyên tắc nền tảng là:
Claim Strength ≤ Evidence Strength.
Vì vậy:
Statistical Significance ≠ Valid Scientific Inference.
1.6. AI Literacy và vấn đề điều kiện biên
AI Literacy ngày càng được xem là năng lực thiết yếu. Tuy nhiên, việc xem AIL như predictor trực tiếp của research quality có thể quá đơn giản.
Một researcher có AI fluency cao nhưng methodological judgment thấp có thể sử dụng AI để tạo một proposal nhanh hơn, phức tạp hơn và thuyết phục hơn về ngôn ngữ nhưng vẫn methodologically misaligned.
Ngược lại, researcher có MF cao và AIL cao có thể sử dụng AI để:
tạo competing alternatives;
challenge methodological recommendations;
kiểm tra assumptions;
xác minh code;
kiểm chứng sources;
đánh giá sensitivity;
nhận diện overclaiming.
Do đó, AIL được đặt như boundary condition của quá trình chuyển methodological judgment thành research architecture.
1.7. Khoảng trống nghiên cứu
Bài viết xác định năm khoảng trống.
Knowledge–Judgment Gap: biết methods chưa đồng nghĩa lựa chọn đúng methods.
Alignment Gap: methodological fit thường được mô tả như thuộc tính của research project nhưng ít được operationalize như researcher competence.
Design–Inference Gap: sampling, measurement, design, analysis và inference thường bị phân mảnh trong đào tạo.
AI Boundary-Condition Gap: AI Literacy thường được nghiên cứu như predictor thay vì moderator của methodological performance.
Measurement Gap: methodological judgment thường được đo bằng self-report thay vì demonstrated performance.
Năm gaps này tạo thành một vấn đề thống nhất:
Làm thế nào methodological knowledge được chuyển thành scientific inference có giá trị trong môi trường nghiên cứu ngày càng có AI hỗ trợ?
- METHODOLOGICAL KNOWLEDGE
Methodological Knowledge (MK) là repertoire kiến thức cần thiết để researcher nhận diện các phương án phương pháp và hiểu assumptions, strengths, limitations và inferential implications của chúng.
MK bao gồm ít nhất:
research design;
sampling;
measurement;
validity;
statistical assumptions;
analytical models;
uncertainty;
causal versus associational inference.
MK là điều kiện cần vì judgment không thể vận hành trong khoảng trống kiến thức.
Tuy nhiên, knowledge không tự chuyển thành fit.
Một researcher có thể biết RCT là thiết kế mạnh cho causal inference nhưng không biết phải làm gì khi randomization không khả thi. Khi đó, methodological competence yêu cầu cân nhắc quasi-experimental alternatives, assumptions và phạm vi causal claim.
Vì vậy:
Methodological Knowledge ≠ Methodological Fit.
- METHODOLOGICAL FIT
3.1. Định nghĩa
Methodological Fit (MF) được định nghĩa là năng lực lựa chọn, liên kết và biện minh sự nhất quán giữa research question, intended inference, construct, measurement, target population, sampling mechanism, research design, data structure, analytical strategy và scientific claim.
MF là relational competence.
Không có technique nào “phù hợp” một cách tuyệt đối. Sự phù hợp chỉ tồn tại trong quan hệ giữa method và problem.
3.2. Question–Design Fit
Research question xác định loại evidence cần thiết.
Descriptive questions, associational questions, predictive questions và causal questions tạo ra những yêu cầu thiết kế khác nhau.
Một causal question được trả lời bằng cross-sectional convenience data có thể tạo association nhưng không đủ để bảo vệ strong causal claim.
Vì vậy:
Research questions specify evidential requirements.
3.3. Construct–Measurement Fit
Construct không đồng nhất với indicator.
Một measure có internal consistency cao chưa chắc đại diện đầy đủ cho construct.
Measurement fit yêu cầu xem xét:
construct definition;
content representation;
response process;
score structure;
reliability;
validity evidence;
intended interpretation.
Do đó, measurement quality phải được đánh giá dựa trên evidence hỗ trợ interpretation and use, không chỉ dựa vào một reliability coefficient.
3.4. Population–Sampling Fit
Sampling nối observed sample với population.
N lớn không tự tạo representativeness.
Population–sampling fit yêu cầu xác định:
target population;
sampling frame;
selection mechanism;
coverage;
nonresponse;
scope of generalization.
Một convenience sample lớn vẫn có thể cung cấp precise estimate cho một systematically selected subset.
Precision ≠ Representativeness.
3.5. Design–Analysis Fit
Analysis phải phù hợp với data-generating process.
Repeated observations tạo dependence.
Nested data tạo clustering.
Latent constructs tạo measurement uncertainty.
Nonrandom assignment tạo confounding.
Vì vậy, lựa chọn analysis phải dựa trên:
Question + Design + Data Structure + Estimand + Assumptions.
Không dựa trên độ phức tạp hoặc tính thời thượng của technique.
3.6. Evidence–Claim Fit
Evidence–claim fit đặt giới hạn cho scientific language.
Correlation không mặc nhiên biện minh causality.
Nonsignificant results không tự chứng minh absence.
Statistical significance không tự chứng minh practical importance.
Good fit indices không chứng minh một theory là explanation duy nhất.
Nguyên tắc:
Claim Strength ≤ Evidence Strength.
- RESEARCH DESIGN QUALITY
4.1. Định nghĩa
Research Design Quality (RDQ) là chất lượng tổng thể của architecture được xây dựng để tạo evidence phù hợp với research question và intended inference.
Ranh giới quan trọng:
MF = quality of alignment decisions.
RDQ = quality of resulting evidence-generating architecture.
Đây là distinction bắt buộc để tránh tautology.
4.2. Các miền RDQ
RDQ gồm:
Research Architecture Completeness;
Design Appropriateness;
Sampling Rigor;
Measurement Quality;
Analysis Architecture;
Validity and Bias Control;
Feasibility and Reproducibility.
Một study có thể có một số alignment decisions hợp lý nhưng overall architecture vẫn yếu do feasibility, incomplete bias control hoặc poor reproducibility.
4.3. Design precedes analysis
Một nguyên tắc cốt lõi:
Good inference begins before analysis.
Design errors thường không thể được “phân tích cho biến mất”.
Sophisticated downstream statistics không tự sửa upstream methodological misfit.
- SCIENTIFIC INFERENCE QUALITY
Scientific Inference Quality (SIQ) phản ánh mức độ final scientific conclusions được hiệu chỉnh theo evidence, design và uncertainty.
SIQ gồm sáu dimensions:
Evidence–Claim Proportionality;
Uncertainty Recognition;
Causal Restraint;
Generalization Appropriateness;
Alternative Explanation Consideration;
Boundary Recognition.
SIQ vì vậy không phải một statistical index.
Một inference tốt phải trả lời:
What does the evidence support?
For whom?
Under what design?
Under which assumptions?
With what uncertainty?
What remains unknown?
Đây là epistemic calibration.
- AI LITERACY NHƯ BOUNDARY CONDITION
AIL trong nghiên cứu này không được định nghĩa đơn giản bằng AI-use frequency hoặc prompt fluency.
AIL có liên quan tới:
AI understanding;
critical evaluation;
methodological recommendation verification;
source verification;
assumption checking;
code validation;
ethical awareness;
inferential restraint.
Đây là điểm quan trọng sau phản biện.
Nếu AIL được dùng như moderator của MF → RDQ, construct phải đủ gần với mechanism cần giải thích.
Do đó:
Generic AI Use ≠ Research-oriented AI Literacy.
AIL có thể khuếch đại methodological competence, nhưng không tự tạo competence.
Một researcher có MF thấp và AIL cao có thể sử dụng AI để triển khai một phương án sai nhanh hơn.
Một researcher có MF cao và AIL cao có thể sử dụng AI để kiểm tra alternatives và cải thiện architecture.
Vì vậy:
AI Literacy = Conditional Capability Amplifier.
- MÔ HÌNH LÝ THUYẾT
Mô hình đề xuất:
METHODOLOGICAL KNOWLEDGE — MK
↓ H1
METHODOLOGICAL FIT — MF
↓ H2
RESEARCH DESIGN QUALITY — RDQ
↓ H3
SCIENTIFIC INFERENCE QUALITY — SIQ
Đồng thời:
MF → SIQ — H4
MF → RDQ → SIQ — H5
MF × AIL → RDQ — H6
MF × AIL → RDQ → SIQ — H7
Logic:
Knowledge → Alignment → Design → Inference
với:
AI Literacy = Boundary Condition.

- GIẢ THUYẾT NGHIÊN CỨU
H1. Methodological Knowledge có quan hệ thuận chiều với Methodological Fit.
H2. Methodological Fit có quan hệ thuận chiều với Research Design Quality.
H3. Research Design Quality có quan hệ thuận chiều với Scientific Inference Quality.
H4. Methodological Fit có quan hệ thuận chiều với Scientific Inference Quality.
H5. Research Design Quality trung gian hóa mối quan hệ giữa Methodological Fit và Scientific Inference Quality.
H6. AI Literacy điều tiết thuận chiều mối quan hệ giữa Methodological Fit và Research Design Quality, sao cho quan hệ MF → RDQ mạnh hơn khi AIL cao hơn.
H7. Tác động gián tiếp của Methodological Fit lên Scientific Inference Quality thông qua Research Design Quality thay đổi theo mức AI Literacy.
Trong observational implementation, H5 và H7 phải được diễn giải dưới dạng indirect/conditional indirect associations consistent with the proposed mechanism, không phải causal effects.
- THIẾT KẾ NGHIÊN CỨU
Nghiên cứu đề xuất multi-method observational design kết hợp structured measures và performance assessment.
Population mục tiêu:
Postgraduate students undertaking quantitative research training.
Quy trình gồm ba thời điểm:
T1 – Baseline Measurement
MK + AIL + theoretically justified covariates.
T2 – Research Design Performance
Standardized Quantitative Research Proposal Task → MF + RDQ.
T3 – Scientific Inference Performance
Standardized Results Package → SIQ.
Temporal separation giúp theoretical ordering rõ hơn nhưng không tự thiết lập causality.
- ĐO LƯỜNG CÁC CONSTRUCT
10.1. Methodological Knowledge
MK được đo bằng:
conceptual items;
scenario-based decisions;
error-detection tasks;
inference-boundary tasks.
Mục tiêu là đo demonstrated knowledge thay vì perceived methodological confidence.
10.2. Methodological Fit
MF được đánh giá qua sáu dimensions:
MF1 – Question–Design Fit;
MF2 – Construct–Measurement Fit;
MF3 – Population–Sampling Fit;
MF4 – Design–Analysis Fit;
MF5 – Evidence–Claim Fit;
MF6 – Inference–Design Fit.
10.3. Research Design Quality
RDQ gồm:
RDQ1 – Architecture Completeness;
RDQ2 – Design Appropriateness;
RDQ3 – Sampling Rigor;
RDQ4 – Measurement Quality;
RDQ5 – Analysis Architecture;
RDQ6 – Bias and Validity Control;
RDQ7 – Feasibility and Reproducibility.
10.4. Scientific Inference Quality
SIQ gồm:
SIQ1 – Evidence–Claim Proportionality;
SIQ2 – Uncertainty Recognition;
SIQ3 – Causal Restraint;
SIQ4 – Generalization Appropriateness;
SIQ5 – Alternative Explanations;
SIQ6 – Boundary Recognition.
10.5. AI Literacy
AIL nên được đo bằng validated instrument phù hợp population, kết hợp khi khả thi với performance indicators về:
AI recommendation verification;
error detection;
source checking;
assumption checking;
code verification.
Không đồng nhất:
Self-reported AI Confidence
với
Demonstrated Evaluative AI Competence.
- CONSTRUCT–MEASUREMENT MATRIX
| Construct | Bản chất | Nguồn đo chính | Câu hỏi đo lường cốt lõi |
| MK | Repertoire | Knowledge/performance test | Researcher biết gì? |
| MF | Relational judgment | Alignment decisions | Các lựa chọn có phù hợp với nhau không? |
| RDQ | Architecture | Proposal artifact | Hệ thống thiết kế tạo evidence tốt đến đâu? |
| SIQ | Epistemic outcome | Inference task | Claim có tương xứng với evidence không? |
| AIL | Boundary capability | Scale + performance | Researcher có thể đánh giá và kiểm chứng AI đến đâu? |
Nguyên tắc discriminant:
MK ≠ MF ≠ RDQ ≠ SIQ.
Đặc biệt:
MF ≠ RDQ.
Nếu empirical evidence không hỗ trợ distinction này, structural model phải được xem xét lại.
- AI-SUPPORTED RESEARCH PROTOCOL
Nếu AI được cho phép trong performance task, conditions phải được chuẩn hóa.
Cần ghi nhận:
AI system/model;
version;
date;
web access;
file access;
duration;
external-source policy;
privacy conditions.
Participants phải lưu AI-use log gồm:
prompts;
AI recommendations;
recommendations accepted;
recommendations rejected;
verification actions;
final human rationale.
Decision provenance:
AI Suggestion → Human Evaluation → Human Decision → Methodological Rationale → Verification Evidence.
Nguyên tắc:
AI Suggested ≠ Researcher Accepted ≠ Methodologically Appropriate.
- HUMAN METHODOLOGICAL ACCOUNTABILITY
AI có thể đề xuất method nhưng không tự sở hữu epistemic responsibility đối với final research claim.
Human Methodological Accountability gồm ba thành phần:
Decision Authority
Justification Obligation
Verification Responsibility.
Researcher phải có khả năng giải thích:
tại sao method được chọn;
alternative nào đã được xem xét;
assumptions nào phải giữ;
AI recommendation nào đã bị từ chối;
sources nào được kiểm chứng;
claim nào evidence không cho phép.
Do đó:
AI Recommendation ≠ Methodological Justification.
- VALIDATION STRATEGY
14.1. Content validity
Expert panel đánh giá relevance, clarity, representativeness, redundancy và construct boundaries.
14.2. Blinded assessment
Artifacts phải được de-identify.
Raters không biết participant identity, MK score hoặc AIL score.
14.3. Rater calibration
Ít nhất hai raters độc lập được đào tạo bằng:
rubric manual;
anchor cases;
practice scoring;
calibration;
predefined adjudication procedure.
14.4. Inter-rater reliability
Continuous rubric scores có thể đánh giá bằng ICC thích hợp.
Ordinal classifications có thể sử dụng weighted κ.
Correlation đơn thuần giữa hai raters không đủ để chứng minh agreement.
14.5. Measurement validity
Trước structural testing cần đánh giá:
dimensionality;
reliability;
convergent evidence;
discriminant evidence;
cross-loadings;
measurement invariance nếu có group comparison.
Đặc biệt phải kiểm tra MF–RDQ discriminant validity.
- PHÂN TÍCH DỮ LIỆU
Structural architecture:
MK → MF
MF → RDQ
RDQ → SIQ
MF → SIQ
MF × AIL → RDQ
MF → RDQ → SIQ
MF → RDQ → SIQ | AIL
Hypothesis–Estimand–Analysis Matrix
| H | Quan hệ | Estimand | Phân tích |
| H1 | MK → MF | Standardized association | SEM/regression |
| H2 | MF → RDQ | Conditional association | SEM |
| H3 | RDQ → SIQ | Conditional association | SEM |
| H4 | MF → SIQ | Direct conditional association | SEM |
| H5 | MF → RDQ → SIQ | Indirect association | Bootstrap |
| H6 | MF × AIL → RDQ | Interaction | Moderation |
| H7 | MF → RDQ → SIQ | AIL | Conditional indirect association | Moderated mediation |
Cỡ mẫu phải được xác định dựa trên primary estimand bằng a priori power analysis hoặc Monte Carlo simulation phù hợp.
Không sử dụng một quy tắc cứng như “10 observations per parameter”.
- COMPETING MODELS VÀ FALSIFIABILITY
Các mô hình cạnh tranh cần được kiểm tra:
AIL → RDQ
AIL → MF
MK → RDQ
MF → RDQ → SIQ without MF → SIQ
Mục tiêu không phải làm dữ liệu phù hợp model đề xuất bằng mọi giá.
Theory phải có khả năng bị bác bỏ.
Nếu MF và RDQ không phân biệt empirically, model phải được sửa.
Nếu H6 không được hỗ trợ, không được kết luận AIL là capability amplifier đối với MF → RDQ.
Nếu conditional indirect effect không xuất hiện, H7 không được coi là supported chỉ vì H5 và H6 riêng rẽ significant.
- ROBUSTNESS VÀ PREREGISTRATION
Robustness analyses có thể bao gồm:
alternative scoring;
models without controls;
alternative AIL specification;
different rater combinations;
subgroup analyses;
robust estimators;
sensitivity to AI-use intensity.
Preregister:
hypotheses;
primary estimands;
sample-size procedure;
measurement model;
scoring;
exclusion rules;
controls;
mediation;
moderation;
conditional indirect effect;
robustness analyses.
Preregistration phải phân biệt confirmatory với exploratory analyses.
- NGUYÊN TẮC BÁO CÁO KẾT QUẢ
Bản thảo hiện tại là theoretical-development/research-protocol article.
Do đó tuyệt đối không tạo:
sample size thực nghiệm giả;
mean/SD giả;
β giả;
p-value giả;
effect size giả;
CFI/TLI/RMSEA/SRMR giả;
confidence interval giả.
Khi nghiên cứu được triển khai, Results nên theo trình tự:
Participant Flow
→ Descriptive Statistics
→ Measurement Model
→ Inter-rater Reliability
→ Discriminant Validity
→ H1–H4
→ H5
→ H6
→ H7
→ Competing Models
→ Robustness Analyses.
Nguyên tắc:
Measurement Model Before Structural Claims.
- THẢO LUẬN
19.1. Từ knowledge đến judgment
Nếu H1 được hỗ trợ, contribution không đơn giản là “knowledge matters”.
Ý nghĩa quan trọng hơn là methodological knowledge tạo nền tảng cho methodological judgment nhưng hai constructs không đồng nhất.
Nếu MK cao nhưng MF thấp, curriculum có thể đang tạo ra:
Knowers of Methods
thay vì:
Judges of Methods.
19.2. MF như missing mechanism
Nếu H2 được hỗ trợ, research competence nên được hiểu:
Knowledge → Judgment → Architecture
thay vì:
Knowledge → Research Quality.
MF trở thành cơ chế chuyển repertoire thành context-sensitive design decisions.
19.3. RDQ như evidence mechanism
Nếu H3 được hỗ trợ, conclusion quan trọng là scientific inference bắt đầu trước statistical analysis.
Measurement, sampling và design đều là inferential decisions.
Do đó:
Good Scientific Inference Begins Before Statistical Analysis.
19.4. AIL như capability amplifier
Nếu H6 được hỗ trợ, không nên kết luận:
AI Literacy improves research quality.
Kết luận chính xác hơn:
AI Literacy strengthens the translation of Methodological Fit into Research Design Quality.
Đây là một conditional proposition.
19.5. Khi AIL không giúp
Nếu H6 không được hỗ trợ, generic AIL có thể không đủ gần methodological mechanism.
Nếu interaction âm, AI fluency có thể trong một số điều kiện khuếch đại automation reliance.
Tuy nhiên, explanation này chỉ hợp lệ khi có process evidence hỗ trợ.
19.6. Ba failure pathways
Failure Pathway A
High MK → Low MF → Weak RDQ → Weak SIQ.
Failure Pathway B
High MF → Weak Implementation → Weak RDQ.
Failure Pathway C
High RDQ → Overclaimed SIQ.
Các pathways cho thấy research quality có thể thất bại ở nhiều tầng khác nhau.
- ĐÓNG GÓP HỌC THUẬT
Đóng góp 1 – Phân biệt repertoire và judgment
MK không được đồng nhất MF.
Đóng góp 2 – Methodological Fit như relational competence
MF được chuyển từ một thuộc tính của study sang competence có thể đánh giá ở researcher.
Đóng góp 3 – RDQ như mechanism
Judgment → Evidence-Generating Architecture → Inference.
Đóng góp 4 – SIQ như epistemic outcome
Research quality cuối cùng không được xác định bởi significance hoặc sophistication mà bởi evidence-calibrated inference.
Đóng góp 5 – AIL như boundary condition
AIL giải thích when methodological judgment được chuyển thành design hiệu quả hơn.
Đóng góp 6 – Human Methodological Accountability
AI có thể hỗ trợ analytical labor nhưng không thay thế trách nhiệm phương pháp và trách nhiệm suy luận của researcher.
- HÀM Ý GIÁO DỤC
Research-methods education cần chuyển từ:
Tool-First Curriculum
sang:
Alignment-First Curriculum.
Thay vì:
t-test → ANOVA → Regression → SEM,
có thể tổ chức:
Question → Intended Inference → Construct → Measurement → Population → Sampling → Design → Analysis → Evidence → Claim.
Assessment cũng cần chuyển từ:
Can the student run the analysis?
sang:
Can the student justify why this analysis is warranted?
Người học phải được yêu cầu giải thích:
Why this design?
Why this measure?
Why this sample?
Why this analysis?
What assumptions?
What alternatives?
What claims are not warranted?
How was AI advice verified?
- HẠN CHẾ
Thứ nhất, observational design không thiết lập causal mediation.
Thứ hai, MF cần thêm validation khi được operationalize như researcher competence.
Thứ ba, MF và RDQ có nguy cơ construct contamination.
Thứ tư, AI Literacy instruments còn heterogeneous.
Thứ năm, self-report AIL có thể khác demonstrated competence.
Thứ sáu, standardized performance task không hoàn toàn tương đương authentic longitudinal research.
Thứ bảy, methodological judgments có thể phụ thuộc discipline.
Thứ tám, rater preferences có thể tạo systematic variation.
Thứ chín, rapid evolution của AI models tạo temporal external-validity problem.
Thứ mười, AIL có thể liên hệ với digital literacy, research experience và general academic competence.
- HƯỚNG NGHIÊN CỨU TIẾP THEO
Một experimental extension có thể so sánh:
No AI
↓
AI Only
↓
AI + Methodological Fit Matrix
↓
AI + Methodological Fit Matrix + Verification Protocol.
Thiết kế này cho phép tách:
AI effect;
alignment-scaffold effect;
verification effect.
Một longitudinal study có thể sử dụng:
T0: MK + AIL
T1: Methodological alignment training
T2: MF + RDQ assessment
T3: SIQ
T4: Unscaffolded transfer task.
Nếu performance duy trì tại T4, bằng chứng về internalized methodological competence sẽ mạnh hơn.
- KẾT LUẬN
GenAI đang làm giảm nhanh rào cản kỹ thuật của quantitative research. Code generation, statistical procedure recommendation, assumption explanation, visualization và preliminary interpretation ngày càng có thể được hỗ trợ bằng AI.
Nhưng chính sự gia tăng analytical accessibility làm nổi bật một distinction quan trọng:
Khả năng thực hiện analysis không đồng nghĩa khả năng biện minh analysis.
Research competence trong kỷ nguyên AI vì vậy không thể chỉ được định nghĩa bằng số lượng techniques researcher biết hoặc tốc độ họ tạo statistical output.
Bài viết đề xuất chuỗi:
Methodological Knowledge
↓
Methodological Fit
↓
Research Design Quality
↓
Scientific Inference Quality
với:
AI Literacy = Boundary Condition của MF → RDQ.
MK cung cấp repertoire.
MF chuyển repertoire thành context-sensitive methodological judgment.
RDQ chuyển judgment thành evidence-generating architecture.
SIQ phản ánh mức độ final scientific claims tương xứng với evidence.
AIL có thể tăng khả năng sử dụng AI để tìm alternatives, kiểm tra assumptions, xác minh code và challenge recommendations, nhưng không tự tạo methodological competence.
Vì vậy, ba mệnh đề trung tâm của nghiên cứu là:
Sophisticated Analysis ≠ Good Research Design.
AI Recommendation ≠ Methodological Justification.
Statistical Significance ≠ Valid Scientific Inference.
Ba mệnh đề tạo thành một architecture thống nhất.
Mệnh đề thứ nhất đặt giới hạn cho technical sophistication.
Mệnh đề thứ hai đặt giới hạn cho AI authority.
Mệnh đề thứ ba đặt giới hạn cho statistical evidence.
Tất cả hội tụ vào một nguyên tắc:
Scientific inference có giá trị đòi hỏi methodological alignment, evidence-generating design và human methodological accountability.
Trong môi trường AI, researcher không cần cạnh tranh với máy về tốc độ tạo code. Năng lực có giá trị hơn là khả năng xác định câu hỏi đúng, lựa chọn evidence phù hợp, nhận diện assumptions, challenge AI recommendations, kiểm chứng output và giới hạn claim.
Do đó, nguyên tắc cuối cùng của bài viết là:
AI may make sophisticated analysis easier; it does not make methodological fit automatic.
Và câu hỏi trung tâm của nghiên cứu định lượng không phải:
“Phương pháp nào phức tạp nhất mà chúng ta có thể sử dụng?”
mà là:
“Phương pháp nào tạo ra loại bằng chứng cần thiết để trả lời câu hỏi này, và bằng chứng đó thực sự cho phép chúng ta kết luận điều gì?”
Đó là điểm methodological knowledge trở thành methodological judgment, judgment trở thành research design, và research design trở thành valid scientific inference.
TUYÊN BỐ
Tuyên bố đóng góp của tác giả theo CRediT
Nguyễn Văn Hùng: Xây dựng ý tưởng và khung khái niệm (Conceptualization); xây dựng phương pháp nghiên cứu (Methodology); phát triển khung lý thuyết (Theoretical Development); thiết kế nghiên cứu (Investigation Design); viết bản thảo ban đầu (Writing – Original Draft); rà soát, phản biện và hiệu chỉnh bản thảo (Writing – Review & Editing).
TÀI LIỆU THAM KHẢO
- Aguinis, H., & Vandenberg, R. J. (2014). An ounce of prevention is worth a pound of cure: Improving research quality before data collection. Annual Review of Organizational Psychology and Organizational Behavior, 1, 569–595. https://doi.org/10.1146/annurev-orgpsych-031413-091231
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
- Creswell, J. W., & Creswell, J. D. (2023). Research design: Qualitative, quantitative, and mixed methods approaches (6th ed.). SAGE Publications.
- Do, H. V., Bui, T. T., Tran, D. H., Nguyen, H. T., & Dinh, L. D. (2026). Developing and validating an AI literacy scale for university students in Vietnam. IFLA Journal, 52(2). https://doi.org/10.1177/03400352251385058
- Edmondson, A. C., & McManus, S. E. (2007). Methodological fit in management field research. Academy of Management Review, 32(4), 1155–1179. https://doi.org/10.5465/AMR.2007.26586086
- Hayes, A. F. (2022). Introduction to mediation, moderation, and conditional process analysis: A regression-based approach (3rd ed.). Guilford Press.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), Article 33267. https://doi.org/10.1525/collabra.33267
- Laupichler, M. C., Aster, A., Schirch, J., & Raupach, T. (2022). Artificial intelligence literacy in higher and adult education: A scoping literature review. Computers and Education: Artificial Intelligence, 3, Article 100101. https://doi.org/10.1016/j.caeai.2022.100101
- Laupichler, M. C., Aster, A., & Raupach, T. (2023). Delphi study for the development and preliminary validation of an item set for the assessment of non-experts’ AI literacy. Computers and Education: Artificial Intelligence, 4, Article 100126. https://doi.org/10.1016/j.caeai.2023.100126
- Lintner, T. (2024). A systematic review of AI literacy scales. npj Science of Learning, 9, Article 50. https://doi.org/10.1038/s41539-024-00264-4
- Miao, F., Shiohira, K., & Lao, N. (2024). AI competency framework for students. UNESCO.
- Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879–903. https://doi.org/10.1037/0021-9010.88.5.879
- Schwarz, J. (2025). The use of generative AI in statistical data analysis and its impact on teaching statistics at universities of applied sciences. Teaching Statistics, 47(2), 118–128. https://doi.org/10.1111/test.12398
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. https://doi.org/10.1080/00031305.2016.1154108
