Reliability Across Borders

Official affiliation for Fernando Parrado
Universidad Sergio Arboleda, Bogotá, Colombia

Corresponding author: fparrado@globalminds.co
ORCID: https://orcid.org/0009-0000-1833-6002

When Reliability Travels Across Borders: Evidence from GLOBE Cultural Dimensions in the United States and Colombia

This is a Working Paper

ABSTRACT

Cross-cultural research depends heavily on the assumption that measurement instruments retain their psychometric properties across countries, languages, and cultural contexts. Although international business scholars have devoted considerable attention to construct development and instrument translation, relatively little research has examined how reliability itself changes when established cultural scales are applied across different national settings.

This working paper investigates the cross-cultural reliability of selected cultural dimensions from the Global Leadership and Organizational Behavior Effectiveness (GLOBE) framework using evidence collected during both a pilot study and a doctoral dissertation comparing negotiators from the United States and Colombia. The analysis reveals that several cultural dimensions exhibited substantial differences in internal consistency across countries, even after careful translation procedures and instrument refinement. Most notably, the In-Group Collectivism construct demonstrated acceptable reliability in Colombia but failed to reach acceptable reliability levels in the United States during the final study, ultimately requiring its removal from the U.S. regression model.

Rather than interpreting these differences solely as methodological shortcomings, this paper argues that variations in reliability may themselves provide meaningful evidence regarding construct equivalence, linguistic interpretation, and cultural conceptualization. The consistency of these findings across both the pilot and final dissertation suggests that researchers should reconsider the widespread assumption that validated cultural instruments automatically preserve their psychometric properties when transferred across national contexts.

The paper contributes to the literature by proposing a conceptual framework that positions reliability differences as potential theoretical evidence rather than merely statistical problems. It also outlines implications for cross-cultural measurement, international negotiation research, and future applications of GLOBE dimensions within International Business studies.

Keywords: Cross-cultural measurement; GLOBE Project; Reliability; Cronbach’s Alpha; Construct Equivalence; International Negotiation; International Business; Translation; Colombia; United States.

  1. INTRODUCTION

During the last four decades, International Business research has increasingly relied upon standardized psychometric instruments to explain differences in managerial behavior across national cultures. Among the most influential frameworks are Hofstede’s cultural dimensions (Hofstede, 1980; Hofstede et al., 2010), the Global Leadership and Organizational Behavior Effectiveness (GLOBE) Project (House et al., 2004), Schwartz’s cultural values, and several complementary approaches that seek to quantify cultural characteristics using standardized survey instruments.

The growing acceptance of these frameworks has significantly improved the comparability of international studies. Researchers working in leadership, negotiation, organizational behavior, international marketing, entrepreneurship, and strategic management routinely adopt validated cultural scales developed elsewhere, translate them into local languages, and assume that the resulting measurements remain conceptually equivalent across countries.

This assumption is fundamental to contemporary cross-cultural research.

If cultural constructs are measured consistently across countries, researchers may legitimately compare means, estimate regression models, examine causal relationships, and propose theoretical explanations regarding the influence of culture on organizational outcomes.

However, this assumption deserves closer examination.

Cross-cultural measurement involves far more than translating words from one language into another. Translation seeks semantic equivalence, yet semantic equivalence does not necessarily imply conceptual equivalence. Individuals raised in different institutional, linguistic, educational, and historical environments may interpret identical survey items differently even when literal translations are accurate.

Consequently, an instrument that performs well in one country may not necessarily preserve its psychometric properties when administered elsewhere.

This concern has existed within the methodological literature for decades. Davis, Douglas, and Silk (1981) warned that measurement reliability represents one of the hidden threats in cross-national research. Marin and Van de Vijver similarly argued that literal translation alone cannot ensure cultural equivalence. Graham and Mintu (1997), studying business negotiations across four countries, observed that translated items should always be considered approximations rather than exact conceptual replicas and suggested that differences in reliability may reflect deeper conceptual variation rather than simple translation errors.

Despite these warnings, relatively few studies explicitly investigate what happens when the reliability of validated cultural scales changes across countries.

Instead, researchers frequently report Cronbach’s Alpha values as routine descriptive statistics before proceeding with substantive hypothesis testing.

When reliability falls below accepted thresholds, the typical response is to eliminate variables, remove questionnaire items, or acknowledge the limitation briefly before continuing the analysis.

This paper proposes a different interpretation rather than viewing reliability differences exclusively as statistical inconveniences, this study argues that they may constitute meaningful empirical evidence regarding the way cultural constructs travel across linguistic and institutional boundaries.

If a construct consistently loses internal consistency when applied within a different national context—even after careful translation, pilot testing, and instrument refinement—perhaps the issue extends beyond measurement error.

Perhaps the construct itself is not interpreted identically across cultures. This possibility becomes especially relevant when studying culture itself. Unlike objective organizational variables such as revenue, firm size, or export intensity, cultural dimensions represent latent psychological constructs. They cannot be observed directly but instead emerge from multiple survey items intended to capture shared underlying values, beliefs, and behavioral tendencies.

If those underlying meanings vary across societies, researchers should expect differences in psychometric behavior. The present study explores this possibility through evidence obtained during a doctoral research project examining cultural predictors of problem-solving behavior in international negotiations between the United States and Colombia.

The broader dissertation employed several cultural dimensions from the GLOBE framework together with measures of interpersonal trust to explain negotiation behavior. Throughout the research process, however, an unexpected methodological pattern emerged.

Reliability coefficients varied considerably across countries.

More importantly, similar reliability concerns appeared during both the pilot study and the final dissertation. The pilot initially suggested that certain GLOBE dimensions exhibited markedly different internal consistency between the two national samples. Rather than disappearing after instrument refinement and a substantially larger final study, several of these patterns persisted, although affecting different constructs. Most notably, In-Group Collectivism ultimately failed to achieve acceptable reliability within the U.S. sample and therefore had to be excluded from the final regression analysis, whereas the same construct demonstrated acceptable reliability among Colombian respondents.

This repeated pattern motivated the present paper.

Instead of treating these findings as methodological anomalies, this working paper examines whether reliability differences themselves deserve theoretical attention within International Business research.

Accordingly, this paper pursues three objectives.

First, it documents the evolution of measurement reliability across both the pilot study and the completed dissertation, providing rare longitudinal evidence regarding the behavior of GLOBE cultural dimensions during instrument refinement.

Second, it reviews the literature on cross-cultural measurement equivalence, translation, and construct validity to position these findings within broader methodological debates.

Finally, it proposes that reliability instability may provide useful evidence regarding conceptual equivalence across cultures and therefore deserves greater attention within International Business research.

Rather than offering definitive answers, this paper should be viewed as the beginning of a broader methodological research agenda. Future research incorporating confirmatory factor analysis, measurement invariance testing, item response theory, and multi-group structural equation modeling will be necessary to determine whether the observed reliability differences originate primarily from translation effects, sampling variation, or genuine conceptual differences between cultures.

Regardless of the ultimate explanation, the evidence presented here suggests that reliability deserves greater attention in cross-cultural research than it has traditionally received. By treating reliability not merely as a statistical prerequisite but also as a potential source of theoretical insight, International Business scholars may improve both the methodological rigor and conceptual richness of comparative cultural research.

  1. Literature Review

2.1 Cross-Cultural Measurement In International Business

The rapid globalization of business has transformed cross-cultural research into one of the central pillars of International Business scholarship. Studies examining leadership, negotiation, international marketing, organizational behavior, entrepreneurship, strategic management, and multinational corporations increasingly depend on standardized psychometric instruments capable of comparing individuals across different national cultures.

Frameworks such as Hofstede’s cultural dimensions, Schwartz’s cultural values, Trompenaars’ cultural orientations, and particularly the Global Leadership and Organizational Behavior Effectiveness (GLOBE) Project have become foundational tools for understanding how culture influences managerial behavior.

The attractiveness of these frameworks lies in their promise of comparability. Researchers assume that individuals answering equivalent questionnaires in different countries are responding to the same latent constructs. Consequently, observed differences may legitimately be interpreted as cultural differences rather than measurement artifacts.

However, this assumption represents one of the greatest methodological challenges facing comparative research. As international business has become increasingly quantitative, researchers have devoted substantial effort to improving sampling procedures, statistical modeling, and analytical sophistication. Comparatively less attention has been devoted to a more fundamental question:

Are identical survey instruments truly measuring the same construct across different cultures?

This question becomes especially important when the constructs themselves represent cultural values.

Unlike objective variables such as firm size or export intensity, cultural dimensions exist only as latent constructs inferred through responses to multiple questionnaire items. Their validity therefore depends upon respondents assigning similar meanings to those items across countries.

When such equivalence does not exist, comparisons become problematic.

2.2 Reliability As More Than A Statistical Requirement

Within psychometric research, reliability is traditionally defined as the degree to which a measurement instrument consistently measures the same underlying construct.

Cronbach’s Alpha remains one of the most widely reported indicators of internal consistency.

In International Business research, reliability coefficients typically appear as preliminary diagnostics before substantive analyses such as regression, structural equation modeling, or multilevel analysis.

If alpha exceeds accepted thresholds, researchers proceed.

If not, variables are often removed.

Although this practice has become routine, it tends to reduce reliability to a technical prerequisite rather than recognizing its potential theoretical implications.

Reliability coefficients are generally interpreted as indicators of measurement quality.

They are seldom interpreted as evidence about the construct itself.

This paper proposes that such an interpretation may be incomplete.

When identical constructs repeatedly display different reliability across countries, researchers should ask whether the phenomenon reflects more than random measurement error.

The observed instability may indicate that respondents organize the underlying concept differently.

Consequently, reliability differences may provide valuable information regarding conceptual equivalence across cultures.

2.3 Translation Does Not Guarantee Conceptual Equivalence

One of the earliest concerns within cross-cultural methodology involved the translation of research instruments.

Brislin’s (1970) back-translation procedure rapidly became the standard approach for adapting questionnaires across languages.

The logic is straightforward.

A questionnaire translated into a second language is independently translated back into the original language.

Differences are discussed until semantic equivalence is achieved.

Although back translation substantially improves linguistic consistency, methodological scholars have long recognized its limitations.

Marin and Marin (1991) argue that literal translation cannot ensure conceptual equivalence because words frequently carry different connotations across cultures.

Similarly, Van de Vijver and Leung distinguish between linguistic equivalence, conceptual equivalence, and measurement equivalence, emphasizing that successful translation addresses only one component of the broader comparability problem.

Consequently, two respondents may answer identical translated questions while activating different cognitive representations.

This issue becomes particularly important for constructs involving values, beliefs, trust, collectivism, leadership, or negotiation behaviors.

Unlike observable phenomena, these concepts depend heavily upon social interpretation.

2.4 Graham And Mintu (1997): An Underappreciated Methodological Contribution

Although Graham and Mintu (1997) are frequently cited within negotiation research because of their behavioral model, one of the most valuable aspects of their study lies in its methodological discussion.

Working with negotiation data collected in the United States, Japan, Brazil, and Spain, the authors confronted substantial reliability differences across countries despite careful translation procedures.

Rather than dismissing these inconsistencies as statistical noise, they explicitly questioned the assumption that translated measures necessarily preserve identical psychometric properties.

The authors observed that translated items are “never more than approximations of the original items” and argued that poor reliability may originate from translation difficulties, conceptual nonequivalence, or genuine cultural differences rather than simple measurement error.

Perhaps even more importantly, Graham and Mintu proposed that formative measurement approaches may sometimes provide a more appropriate representation of culturally complex constructs than traditional reflective measurement models.

Although their empirical context involved negotiation behaviors rather than GLOBE dimensions, the methodological implications extend well beyond negotiation research.

The present study follows this methodological tradition.

However, instead of comparing four countries simultaneously, it examines in greater detail the evolution of measurement reliability during both the pilot and final stages of a single cross-cultural research project comparing the United States and Colombia.

2.5 The GLOBE Project and Measurement Challenges

The Global Leadership and Organizational Behavior Effectiveness (GLOBE) Project represents one of the most ambitious efforts ever undertaken to understand cultural influences on leadership and organizational behavior.

Building upon Hofstede’s pioneering work, GLOBE expanded the conceptualization of national culture by incorporating nine cultural dimensions measured across dozens of societies.

Its influence within International Business has been considerable.

Researchers have applied GLOBE dimensions to leadership, innovation, strategic management, entrepreneurship, human resource management, international alliances, and negotiation.

Because of its rigorous development process, the framework is generally considered highly reliable.

Nevertheless, GLOBE was originally designed to capture broad societal patterns rather than specific behavioral domains such as negotiation.

Consequently, applying GLOBE dimensions within different research contexts requires careful consideration of construct validity.

Furthermore, although the GLOBE project devoted enormous effort to translation and cross-cultural coordination, subsequent applications frequently adapt subsets of dimensions, different respondent populations, and alternative research contexts.

Each adaptation potentially introduces additional sources of measurement variation.

Therefore, reliability should not be assumed.

It should be empirically demonstrated.

2.6 Measurement Equivalence: The Next Frontier

Contemporary methodological literature increasingly emphasizes that measurement equivalence represents a prerequisite for meaningful cross-cultural comparisons.

Steenkamp and Baumgartner, Vandenberg and Lance, Davidov and colleagues, and Van de Vijver and Leung argue that establishing equivalence requires evidence at multiple levels.

Researchers should evaluate:

configural equivalence, metric equivalence, scalar equivalence, functional equivalence, and conceptual equivalence.

Cronbach’s Alpha alone cannot establish measurement equivalence.

Nevertheless, it frequently provides the first indication that equivalence may be problematic.

When substantial differences in internal consistency repeatedly emerge across national samples, they signal the need for deeper psychometric investigation.

The present paper does not claim to resolve these issues definitively.

Instead, it proposes that reliability differences deserve greater theoretical attention because they may identify constructs requiring further examination through confirmatory factor analysis, measurement invariance testing, and multi-group structural equation modeling.

2.7 From A Methodological Problem To A Research Opportunity

Most cross-cultural studies report reliability coefficients only briefly before proceeding to hypothesis testing.

Variables exhibiting low reliability are often eliminated with little further discussion.

This practice may inadvertently overlook valuable theoretical information.

The evidence presented in this paper suggests an alternative interpretation.

Repeated reliability differences observed during both the pilot study and the completed dissertation indicate that some cultural dimensions may not travel across national borders as seamlessly as commonly assumed.

Rather than viewing these differences exclusively as methodological limitations, they may be interpreted as empirical clues regarding how cultural constructs evolve when translated across languages, institutions, and national contexts.

This perspective transforms reliability from a statistical prerequisite into a research question.

Instead of asking only whether an instrument is reliable, researchers might also ask:

What Does It Mean When A Cultural Construct Systematically Changes Its Psychometric Behavior Across Countries?

That question forms the central motivation for the empirical analysis presented in the following sections.

  1. Research Context, Methodology And Empirical Evidence

3.1 Research Context

The empirical evidence presented in this working paper originates from a broader doctoral research project examining the influence of national culture on international negotiation behavior between the United States and Colombia. The dissertation sought to extend the Problem-Solving Approach (PSA) negotiation model proposed by Graham and colleagues by incorporating cultural dimensions from the Global Leadership and Organizational Behavior Effectiveness (GLOBE) framework together with interpersonal trust as explanatory variables.

The broader research objective was substantive rather than methodological. Specifically, the dissertation investigated whether selected cultural dimensions predicted negotiators’ tendency to adopt a collaborative problem-solving approach during international business negotiations.

However, throughout the research process, an unexpected methodological phenomenon emerged.

Several cultural dimensions exhibited markedly different reliability levels across countries.

Initially, these differences appeared during the pilot study.

After instrument refinement, questionnaire revision, additional data collection, and completion of the final dissertation, similar reliability differences continued to emerge.

Rather than disappearing through improved research design, measurement instability remained a persistent characteristic of several constructs.

This observation motivated the present methodological analysis.

Instead of focusing on negotiation outcomes, this working paper investigates the behavior of the measurement instrument itself.

3.2 Research Design

The original dissertation employed a quantitative cross-sectional design involving participants from two countries:

  • United States
  • Colombia

The study used established psychometric scales rather than developing new measurement instruments.

The principal independent variables were selected GLOBE cultural dimensions:

  • Institutional Collectivism
  • Future Orientation
  • Power Distance
  • In-Group Collectivism
  • Uncertainty Avoidance
  • Gender Egalitarianism
  • Performance Orientation

In addition, interpersonal Trust was incorporated as an independent construct because previous literature suggests that trust plays an important role during collaborative negotiations.

The dependent variable was the negotiator’s Problem-Solving Approach (PSA).

Because all cultural dimensions represented latent psychological constructs measured through multiple questionnaire items, internal consistency became an essential prerequisite before hypothesis testing.

Accordingly, Cronbach’s Alpha coefficients were calculated independently for each national sample.

3.3 The Pilot Study: Early Warning Signals

Like many international research projects, the dissertation began with a pilot study designed primarily to evaluate questionnaire clarity, translation quality, respondent understanding, and preliminary psychometric behavior.

The pilot fulfilled these traditional objectives.

However, it also revealed an unexpected pattern.

Several cultural dimensions displayed substantially different reliability coefficients between the United States and Colombia.

Table 1 summarizes the pilot reliability results.

Table 1.

Cronbach’s Alpha – Pilot Study

[Aquí irá la tabla del piloto de GLOBE.]

The pilot results suggested that measurement instability was not random.

Rather, specific constructs appeared sensitive to national context.

For example, Institutional Collectivism produced relatively weak reliability in the United States (α = .473) while achieving substantially stronger internal consistency in Colombia (α = .779).

Conversely, In-Group Collectivism behaved in the opposite direction, reaching acceptable reliability in the United States (α = .665) but exhibiting very low internal consistency in Colombia (α = .336).

Similarly, Uncertainty Avoidance, Gender Egalitarianism, and Performance Orientation produced noticeably weaker reliability among Colombian respondents than among participants from the United States.

At this stage, several possible explanations were considered.

First, differences could reflect normal sampling variability associated with pilot studies.

Second, translation issues might have affected respondents’ interpretation of specific questionnaire items.

Third, respondents in both countries might simply conceptualize some cultural dimensions differently.

Because pilot studies frequently reveal temporary psychometric issues that disappear after instrument refinement, no definitive conclusions were drawn.

Instead, the questionnaire was carefully reviewed before launching the final dissertation.

3.4 Instrument Refinement

Following the pilot study, several improvements were implemented.

Question wording was reviewed to improve clarity.

Translation consistency was reexamined.

Survey administration procedures were standardized.

The final dissertation also incorporated a substantially larger and more representative sample.

Under conventional psychometric expectations, one would anticipate improved measurement stability.

Surprisingly, although several constructs improved, important reliability differences remained.

Even more interestingly, the specific constructs exhibiting lower reliability changed between the pilot and the final study.

This observation suggests that the phenomenon cannot easily be attributed to isolated translation errors.

Instead, it points toward a more complex interaction between language, culture, and construct interpretation.

3.5 Evidence from the Final Dissertation

The completed dissertation confirmed that several GLOBE dimensions demonstrated satisfactory internal consistency across both countries.

Trust, for example, exhibited excellent reliability in both national samples.

Gender Egalitarianism and Power Distance also performed consistently well.

However, other dimensions continued to display important differences.

Table 2 summarizes the final reliability coefficients.

Table 2.

Cronbach’s Alpha – Final Dissertation

The most striking finding concerns In-Group Collectivism.

Whereas Colombian respondents produced an acceptable Cronbach’s Alpha of approximately .774, the corresponding reliability coefficient for the United States reached only .408.

Because this value fell well below accepted psychometric standards, the construct was removed from the final regression model for the U.S. sample.

Future Orientation also exhibited weaker reliability in the United States (α = .598) than in Colombia (α = .724), although the difference was less dramatic.

In contrast, Trust maintained exceptionally high reliability across both countries (United States α = .946; Colombia α = .888), suggesting that not all constructs experienced cross-cultural instability.

This contrast becomes particularly important.

If translation alone explained the observed differences, similar reliability problems would be expected across all constructs.

Instead, only specific cultural dimensions consistently demonstrated instability.

Such selective behavior suggests that conceptual interpretation rather than simple linguistic translation may play an important role.

3.6 Comparing the Pilot and Final Studies

The comparison between the pilot and final dissertation provides perhaps the most valuable contribution of this working paper.

Rather than presenting a single reliability assessment, the research documents how measurement properties evolved during the development of an international research project.

Several observations emerge.

First, reliability differences were already present during the pilot.

Second, they did not disappear after questionnaire refinement.

Third, the specific constructs exhibiting instability changed over time.

Finally, some constructs remained consistently robust across both countries.

Taken together, these findings suggest that measurement reliability in cross-cultural research should be viewed as a dynamic characteristic rather than a fixed property of an instrument.

The results also caution against interpreting low reliability exclusively as evidence of poor questionnaire design.

Instead, repeated differences across countries may indicate that respondents organize certain cultural concepts differently.

This possibility deserves considerably greater methodological attention than it has traditionally received within International Business research.

3.7 A Different Interpretation of Reliability

The conventional interpretation of Cronbach’s Alpha is straightforward:

higher values indicate better measurement.

Lower values indicate weaker internal consistency.

Although statistically correct, this interpretation may be incomplete in cross-cultural research.

When the same construct repeatedly demonstrates different reliability across countries despite careful translation, pilot testing, and instrument refinement, the question shifts from:

“Is this instrument reliable?”

to a more theoretically interesting one:

“Is this construct being understood in the same way across cultures?”

The evidence presented here cannot answer that question definitively.

However, it strongly suggests that reliability differences themselves deserve further investigation through more sophisticated psychometric techniques such as confirmatory factor analysis, measurement invariance testing, item response theory, and multigroup structural equation modeling.

  1. Discussion

4.1 Reliability as a Source of Theoretical Insight

The principal argument developed throughout this paper is that reliability should not be viewed exclusively as a statistical prerequisite for hypothesis testing. Within cross-cultural research, differences in internal consistency may themselves provide valuable theoretical information regarding construct equivalence.

The conventional research process generally follows a familiar sequence. Researchers translate an existing instrument, estimate reliability coefficients, eliminate variables that fail to satisfy accepted thresholds, and proceed with substantive statistical analyses. Low reliability is therefore treated as a methodological limitation that must be corrected before the “real” research begins.

The findings presented in this working paper suggest that such an interpretation may overlook important information.

Throughout both the pilot study and the final dissertation, several GLOBE cultural dimensions exhibited systematic differences in internal consistency across the United States and Colombia. Importantly, these differences persisted despite improvements in questionnaire design, translation review, sample expansion, and instrument refinement.

If measurement problems remain after these methodological improvements, researchers should consider the possibility that reliability differences reflect more than technical deficiencies.

They may indicate that respondents from different cultural environments organize the underlying construct differently.

Such an interpretation is consistent with contemporary discussions regarding conceptual equivalence within comparative research.

4.2 Beyond Translation: Conceptual Equivalence

Translation represents only the first step toward achieving comparable measurement across cultures.

Back translation procedures substantially improve semantic equivalence, yet semantic equivalence does not necessarily imply conceptual equivalence.

Words may be translated accurately while still activating different cognitive schemas among respondents from different societies.

This possibility appears particularly relevant for cultural constructs.

Concepts such as collectivism, trust, uncertainty avoidance, future orientation, and gender egalitarianism emerge from broader social institutions rather than from isolated individual experiences.

Consequently, respondents may assign different meanings to apparently identical questionnaire items.

If this interpretation is correct, then differences in reliability should not automatically be interpreted as evidence of poor questionnaire quality.

Instead, they may provide indirect evidence that the construct itself requires additional conceptual clarification before meaningful international comparison becomes possible.

4.3 Implications for the GLOBE Framework

The present paper does not question the validity of the GLOBE Project.

On the contrary, the GLOBE framework remains one of the most comprehensive and influential cultural models available within International Business.

However, the present findings suggest that subsequent applications of GLOBE dimensions deserve careful psychometric evaluation whenever they are transferred to new national contexts, different respondent populations, or alternative research domains.

Researchers frequently adopt GLOBE dimensions assuming that published reliability estimates automatically generalize to their own samples.

The evidence presented here suggests that such assumptions should be verified empirically rather than accepted automatically.

Reliability should become an explicit research finding rather than a routine descriptive statistic.

4.4 Implications for International Business Research

The methodological implications extend beyond negotiation research.

International Business increasingly depends upon multinational datasets involving multiple languages, institutional environments, and cultural settings.

Comparative studies investigating leadership, entrepreneurship, innovation, organizational commitment, strategic alliances, corporate governance, and international management frequently rely upon psychometric scales originally developed elsewhere.

The findings reported here suggest that greater attention should be devoted to measurement equivalence before substantive cross-national comparisons are interpreted.

Future studies should routinely complement reliability analyses with confirmatory factor analysis, measurement invariance testing, and multigroup structural equation modeling whenever possible.

Doing so will improve both methodological rigor and theoretical confidence.

  1. LIMITATIONS

This working paper represents an initial methodological reflection rather than a definitive psychometric investigation.

Several limitations should therefore be acknowledged.

First, the analysis relies on evidence obtained from a single doctoral research project involving two countries. Although the comparison between the pilot study and the completed dissertation provides valuable longitudinal insight, broader multinational studies would strengthen the generalizability of the findings.

Second, the present analysis focuses primarily on internal consistency measured through Cronbach’s Alpha. Future research should incorporate additional indicators such as Composite Reliability, McDonald’s Omega, Average Variance Extracted, and Confirmatory Factor Analysis.

Third, this paper does not formally test measurement invariance. Consequently, no definitive conclusions regarding configural, metric, or scalar equivalence can yet be drawn.

Finally, because the original dissertation was designed to investigate negotiation behavior rather than psychometric equivalence, some methodological questions identified here remain open for future investigation.

  1. FUTURE RESEARCH

This working paper opens several promising avenues for future research.

First, future studies should investigate whether the reliability differences identified here persist across additional countries using the same GLOBE dimensions.

Second, multigroup Confirmatory Factor Analysis should be employed to evaluate configural, metric, scalar, and residual invariance.

Third, Item Response Theory could identify specific questionnaire items contributing disproportionately to reliability differences across countries.

Fourth, qualitative research involving cognitive interviews may help explain how respondents from different cultures interpret identical questionnaire items.

Finally, similar analyses should be extended beyond GLOBE to other widely used international instruments, including Hofstede, Schwartz, and World Values Survey measures.

Such research would substantially improve the methodological foundations of comparative International Business research.

  1. CONCLUSION

Cross-cultural research depends fundamentally upon the assumption that measurement instruments preserve their conceptual meaning across countries.

The evidence presented in this working paper suggests that this assumption deserves closer examination.

Using data collected during both a pilot study and a completed doctoral dissertation comparing the United States and Colombia, this paper documented systematic reliability differences across several GLOBE cultural dimensions.

Rather than interpreting these findings exclusively as methodological shortcomings, this paper proposed an alternative perspective.

Reliability differences may themselves provide valuable evidence regarding conceptual equivalence across cultures.

Consequently, psychometric instability should not always be viewed as a problem to eliminate.

Sometimes, it may represent an opportunity to better understand how culture shapes the very constructs researchers seek to measure.

By treating reliability as both a methodological and theoretical phenomenon, International Business scholars may develop more robust instruments, stronger comparative theories, and richer explanations of cultural behavior across national contexts.

References For The Next Version (Apa – To Expand)

Bibliografía base para convertir este working paper en un artículo publicable.

  1. Cross-Cultural Measurement
  • Brislin, R. W. (1970). Back-Translation for Cross-Cultural Research.
  • Davis, H., Douglas, S., & Silk, A. (1981). Measurement Reliability: A Hidden Threat to Cross-National Marketing Research.
  • Van de Vijver, F., & Leung, K. (1997). Methods and Data Analysis for Cross-Cultural Research.
  • Van de Vijver, F., & Leung, K. (2021). Measurement Equivalence in Cross-Cultural Research.
  • Harkness, J. A. (2003). Questionnaire Translation.
  • Harkness, J., Van de Vijver, F., & Mohler, P. (2003). Cross-Cultural Survey Methods.
  • Marin, G., & Marin, B. V. (1991). Research with Hispanic Populations.
  • Babbie, E. (2016). The Practice of Social Research.
  1. Measurement Invariance
  • Steenkamp, J.-B. E. M., & Baumgartner, H. (1998). Assessing Measurement Invariance in Cross-National Consumer Research.
  • Vandenberg, R., & Lance, C. (2000). A Review and Synthesis of the Measurement Invariance Literature.
  • Cheung, G., & Rensvold, R. (2002). Evaluating Goodness-of-Fit Indexes for Testing Measurement Invariance.
  • Davidov, E., Meuleman, B., Cieciuch, J., Schmidt, P., & Billiet, J. (2014). Measurement Equivalence in Cross-National Research.
  • Byrne, B. (2016). Structural Equation Modeling with AMOS.
  1. GLOBE and Culture
  • House, R. J., Hanges, P., Javidan, M., Dorfman, P., & Gupta, V. (2004). Culture, Leadership, and Organizations: The GLOBE Study of 62 Societies.
  • Javidan, M., et al. (GLOBE 2020 papers).
  • Hofstede, G. (1980). Culture’s Consequences.
  • Hofstede, G. (2001). Culture’s Consequences (2nd ed.).
  • Hofstede, G., Hofstede, G. J., & Minkov, M. (2010). Cultures and Organizations.
  • Schwartz, S. H. (1994). Beyond Individualism and Collectivism.
  1. Negotiation and Methodology
  • Graham, J. L. (1986). The Problem-Solving Approach to Negotiation.
  • Graham, J. L., Mintu-Wimsatt, A., & Rodgers, W. (1994). Negotiation Behaviors in Ten Countries.
  • Graham, J. L., & Mintu-Wimsatt, A. (1997). Culture’s Influence on Business Negotiations in Four Countries.
  • Rubin, J., & Brown, B. (1975). The Social Psychology of Bargaining and Negotiation.
  • Walton, R., & McKersie, R. (1965). A Behavioral Theory of Labor Negotiations.