Abstract
The Berliner Studierfähigkeitstest—Psychologie (BSF-P), known internationally as the Berlin Aptitude Test for Psychology, is a standardized psychometric battery engineered to assess cognitive and scholastic aptitude specifically tailored to undergraduate degree programs in psychology. Developed in Germany to address the acute institutional demands of high-stakes university admissions—where degree programs face rigorous numerical restrictions (numerus clausus)—the BSF-P evaluates study potential through six psychometrically calibrated subtests: Reading Comprehension, English Language Skills, Math Skills, Verbal Reasoning, Numerical Reasoning, and Figural Reasoning. The test was conceptualized and calibrated by Kai T. Horstmann, Andra Biesok, Katja Witte, Henrik R. Godmann, Karla Fliedner, Lisa Wilm, Larissa Doran, and Matthias Ziegler at Humboldt-Universität zu Berlin and the University of Siegen. Test content was derived through an evidence-based blueprint combining structured critical-incident interviews with university professors, researchers, and students alongside systematic reviews of academic performance determinants. Utilizing Item Response Theory (IRT) parameter estimation (specifically Rasch-family scaling), applicant proficiency is estimated via standardized personal parameters ($z$-scores) across subtests and aggregated into a composite aptitude indicator. Psychometric evaluations across multiple piloting cohorts and operational selection cycles demonstrate sound structural validity ($\chi^2(50) = 194.26$, $ ext{CFI} = .899$,$ ext{RMSEA} = .078$,$ ext{SRMR} = .054$), robust composite test battery reliability ($lpha / ext{composite reliability} = .81$ to $.92$), and test-retest reliability across an 84-day interval of $r = .74$ ($r = .51$ to $.65$ across subscales). The battery exhibits discriminant validity against personality dimensions (Big Five Inventory-2) and vocational interests, while demonstrating incremental validity over secondary school grade point averages ($r = .26$, $p < .01$). The BSF-P represents a vital psychometric advance in European higher education selection, standardizing applicant evaluation while mitigating exclusive reliance on high school diplomas.
Keywords
Academic Ability, Academic Aptitude, Bachelor of Psychology, Incremental Validity, Measurement Precision, Standardized Aptitude Test, Test Battery, Undergraduate Psychology, University Applicants, University Student Admissions, University Student Selection, Item Response Theory
Authors
The Berlin Aptitude Test for Psychology was conceptualized, constructed, and psychometrically validated by a specialized research team in psychological diagnostics and differential psychology across Humboldt-Universität zu Berlin and the University of Siegen:
- Kai T. Horstmann — Department Psychologie, Universität Siegen, Siegen, Germany.
- Andra Biesok — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Katja Witte (ORCID: 0009-0007-3718-6928) — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Henrik R. Godmann — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Karla Fliedner — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Lisa Wilm — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Larissa Doran — Institut für Psychologie, Humboldt-Universität zu Berlin, Berlin, Germany.
- Matthias Ziegler (Corresponding Author) — Institut für Psychologie, Psychologische Diagnostik, Humboldt-Universität zu Berlin, Unter den Linden 6, 10099 Berlin, Germany. Email: [email protected].
Purpose
The primary purpose of the Berliner Studierfähigkeitstest—Psychologie (BSF-P) is to provide an objective, standardized, and psychometrically robust instrument for assessing subject-specific scholastic aptitude among applicants seeking admission into Bachelor of Science (B.Sc.) psychology programs. In Germany and across numerous European higher education systems, undergraduate psychology programs represent some of the most competitive academic courses. Selection has historically relied almost exclusively on secondary school graduation GPAs (the German Abitur). However, exclusive reliance on high school grades introduces significant challenges, including regional grading disparities, grade inflation, and lack of domain specificity regarding the advanced quantitative, analytical, and linguistic demands inherent to modern empirical psychological science.
The BSF-P was created to resolve these selection challenges by establishing a standardized, legally defensible, and empirically validated admissions battery. Modern psychology curricula require not only foundational verbal and communicative capabilities, but also rigorous quantitative reasoning, inferential statistics, empirical experimental methodology, and comprehension of advanced international scientific literature primarily published in English. The BSF-P explicitly evaluates these multidimensional competencies, allowing university admissions committees to evaluate candidates on cognitive dimensions directly predictive of academic coursework performance.
Beyond institutional student selection, the BSF-P serves critical research and practical functions in educational measurement. In clinical and counseling settings related to educational guidance, the test provides diagnostic insights into an individual’s specific cognitive strengths and deficits relative to the demands of empirical university curricula. For institutional researchers, the battery provides a standardized metric to examine cognitive determinants of higher education success, longitudinal academic persistence, drop-out attrition, and the compensatory interplay between general cognitive ability, personality traits, and prior scholastic achievement.
Psychological Construct
The underlying construct assessed by the BSF-P is Psychology Academic Aptitude (Studierfähigkeit für das Fach Psychologie). Rather than operating as an undifferentiated general intelligence ($g$-factor) test, the BSF-P conceptualizes academic aptitude as a domain-specific composite of fluid intelligence, crystallized knowledge, and operational scholastic skills critical for navigating contemporary scientific psychology. The battery operationalizes this overarching construct through six differentiated subconstructs:
1. Reading Comprehension (Leseverständnis)
This dimension measures the ability to extract, synthesize, evaluate, and critically appraise complex theoretical and methodological arguments presented in dense scientific texts. In psychological education, students must continuously digest intricate empirical articles, identify experimental hypotheses, evaluate methodological designs, and detect logical fallacies or unsupported conclusions. The subtest presents long-form academic passages followed by multi-layered questions evaluating inferential depth rather than superficial verbatim recall.
2. English Language Skills (Englischkenntnisse)
Because contemporary psychological science is conducted and communicated overwhelmingly in English, proficiency in receptive academic English is an indispensable prerequisite for university success. This dimension evaluates the student’s mastery of specialized academic vocabulary, syntax parsing, and contextual comprehension within scientific passages, assessing whether matriculating students can understand international research literature, statistical reporting, and foundational textbooks without substantial language barriers.
3. Math Skills (Mathematische Fähigkeiten)
Psychology programs demand an extensive foundation in mathematics to facilitate mastery of descriptive and inferential statistics, psychometrics, experimental design, and quantitative modeling. This subscale measures operational fluency in algebraic manipulation, probability estimation, fractions, percentages, basic calculus concepts, and geometric transformations. The items focus on applied mathematical problem-solving relevant to data interpretation and statistical computation.
4. Verbal Reasoning (Verbales Schließen)
Reflecting crystallized verbal intelligence and formal deductive capacity, this subscale captures the capacity to discern semantic relationships, comprehend verbal analogies, process syllogistic logic, and derive logically valid deductions from complex linguistic premises. Students are required to distinguish between necessary logical consequences and plausible but non-deductive assertions, a competency vital for analyzing psychological theories and constructing coherent empirical arguments.
5. Numerical Reasoning (Numerisches Schließen)
Representing fluid quantitative intelligence, numerical reasoning involves the capacity to identify abstract rules, detect patterns within numerical matrices or number series, and induce mathematical algorithms governing sequential data. This dimension is central to statistical thinking, enabling students to recognize underlying trends, evaluate probability distributions, and interpret complex data outputs in quantitative psychological research.
6. Figural Reasoning (Figurales Schließen)
Rooted in classical fluid reasoning (Cattell’s $Gf$), figural reasoning evaluates the applicant’s ability to manipulate spatial configurations mentally, identify topological patterns, and solve abstract non-verbal matrix problems. This dimension ensures that the BSF-P captures domain-general inductive reasoning capabilities unconfounded by prior formal education, language proficiency, or cultural background.
Theoretical Framework
The theoretical framework of the BSF-P synthesizes contemporary structural intelligence theories with classical personnel selection and educational measurement models. Foremost among these is the Cattell-Horn-Carroll (CHC) theory of cognitive abilities (Schneider & McGrew, 2018). CHC theory posits a three-stratum hierarchical structure of human cognitive abilities, encompassing narrow abilities (Stratum I), broad cognitive domains (Stratum II), and general cognitive ability (Stratum III). The BSF-P systematically samples broad Stratum II domains directly linked to university performance: Fluid Reasoning ($Gf$, captured via Figural and Numerical Reasoning), Comprehension-Knowledge / Crystallized Intelligence ($Gc$, captured via Verbal Reasoning and Reading Comprehension), Quantitative Knowledge ($Gq$, captured via Math Skills), and Reading and Writing capabilities ($Grw$, captured via English and German Reading Comprehension).
In parallel, the battery incorporates the Berlin Model of Intelligence Structure (BIS; Jäger, 1982, 1984), which emphasizes that cognitive performance is structured bi-factorially by operational modes (e.g., processing capacity, memory, creativity, speed) and content facets (verbal, numerical, and figural). By intentionally configuring its six subtests across verbal, numerical, and figural content facets, the BSF-P ensures a balanced representation of cognitive modalities, minimizing single-facet test bias.
From an applied selection perspective, the BSF-P is grounded in Campbell’s Theory of Job and Academic Performance and the principles of construct-oriented test construction. Rather than sampling generic aptitude in an educational vacuum, Horstmann and colleagues utilized a criterion-centric development approach. By conducting structured critical-incident interviews with academic faculty, researchers, and students alongside meta-analytic synthesis of predictors of academic success, the authors established a concrete nomological network mapping specific cognitive aptitudes to the authentic performance challenges encountered in undergraduate psychology degree programs.
Validity
The BSF-P development team executed a rigorous psychometric program to establish multiple lines of construct, content, discriminant, and incremental validity evidence in alignment with the Standards for Educational and Psychological Testing (AERA, APA, & NCME):
Content Validity
Content validity was established through systematic job and curriculum analyses. Items were authored according to comprehensive evidence-based domain specifications derived from interviews with academic professors, undergraduate instructors, and senior researchers. Independent validation comes from convergence with other established German psychology admission procedures, such as the tests described by Formazin et al. (2011) and the standardized admission test for psychology (STAV-Psych; Watrin et al., 2022). These independent batteries share corresponding subtest structures and operational definitions, demonstrating consensus regarding the foundational content domain of psychological study aptitude.
Discriminant Validity
To verify that the BSF-P measures cognitive aptitude rather than non-cognitive traits or general vocational preferences, the authors administered the battery alongside the German adaptation of the Big Five Inventory-2 (BFI-2; Danner et al., 2016) and standardized measures of academic and vocational interests (the Revised Inventory of Academic Interests, RIAS; Roemer et al., 2021; Rounds et al., 2010). Empirical correlations between BSF-P subtest scores and the Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism) were consistently weak to negligible, confirming that the test does not confound cognitive aptitude with personality style. Similarly, correlations with academic interest dimensions confirmed that the cognitive parameters of the BSF-P are distinct from motivational preferences.
Incremental Validity
A crucial psychometric test of any university admissions instrument is its ability to account for variance beyond traditional school leaving qualifications. The correlation between the German higher education entrance qualification (Abitur grade) and the overall BSF-P test score was established at:
$r = .26$ (95% CI [.18; .35], $p < .01$)
This moderate correlation indicates that while the BSF-P shares a modest common core with cumulative high school achievement, roughly 93% of its variance remains unshared with secondary school GPAs. Consequently, the BSF-P contributes substantial incremental information regarding domain-specific cognitive competencies, offering admissions committees predictive utility beyond traditional academic credentials.
Construct and Structural Validity
The overarching construct validity of the battery is supported by Item Response Theory (IRT) analyses and confirmatory structural equation modeling (SEM). Rasch-family item calibrations confirmed that the items within each subtest scale unidimensionally across test takers. Structural equation models specifying an overarching study aptitude factor accounting for common variance across the six subtests yielded acceptable goodness-of-fit indices across diverse applicant cohorts.
Limitations: Predictive Validity
A notable limitation documented by Horstmann et al. (2023) is that longitudinal predictive validity coefficients linking initial BSF-P scores to cumulative undergraduate grade point averages, exam retake rates, and final thesis evaluations could not be fully reported at initial test publication due to the multi-year cycle required for cohorts to progress through university study. Longitudinal tracking of admitted cohorts at Humboldt University of Berlin remains ongoing to confirm the battery’s predictive forecasting precision.
Reliability
The BSF-P displays high measurement precision across internal consistency indicators and longitudinal temporal stability assessments:
Internal Consistency and Battery Precision
Given the multi-component battery structure of the BSF-P, composite reliability was estimated using standardized procedures for multidimensional test batteries outlined by Bühner (2021). For the composite overall BSF-P score, the battery reliability reached:
- Composite Battery Reliability (Initial Administration): $lpha / r_{ ext{co\mp}} = .92$
- Composite Battery Reliability (Subsequent / Final Operational Cohorts): $lpha / r_{ ext{co\mp}} = .81$
For the individual subtests, internal consistency values (Cronbach’s alpha and model-based reliability estimates) ranged from $lpha = .50$ to $.76$. While individual subtest reliabilities vary—reflecting brief testing lengths designed to minimize applicant fatigue during high-stakes selection—the composite overall score delivers high precision sufficient for individual-level decision making.
Test-Retest Reliability (Temporal Stability)
Temporal stability was evaluated across an average retest interval of 84 days (approximately 12 weeks), representing a realistic timeline between preliminary test preparation and operational admissions cycles:
- Overall BSF-P Composite Score Retest Stability: $r_{tt} = .74$
- Individual Subtests Retest Stability: Ranged from $r_{tt} = .51$ to $r_{tt} = .65$
These stability coefficients confirm that the BSF-P measures enduring cognitive aptitudes rather than transient psychological states, demonstrating stability across testing intervals.
Factor Analysis
The dimensionality and structural architecture of the BSF-P were validated through modern psychometric scaling procedures combining Item Response Theory (IRT) and structural equation modeling (SEM):
Item Response Theory (IRT) Analyses
Individual subtests were evaluated using 1-Parameter Logistic (1PL / Rasch-family) measurement models. These IRT scaling procedures confirmed:
- Unidimensionality and Item Homogeneity: Items within each of the six subtests loaded consistently onto their respective latent trait dimension, satisfying prerequisites for invariant item and person ordering across test-takers (Bühner, 2021).
- Difficulty Parameter Distribution: Item difficulty distributions covered a broad latent continuum, with targeted enrichment in the moderate-to-high difficulty range ($ heta > +1.0$). This intentional calibration ensures high measurement precision and test information in the upper ability spectrum, preventing ceiling effects and maximizing discrimination among high-achieving applicants competing for restricted university placements.
Structural Equation Modeling (SEM)
Latent person parameters ($ heta_i$) estimated via the IRT measurement models for each of the six subtests were incorporated into an overarching structural model representing general psychology academic aptitude. The structural measurement model demonstrated acceptable fit to the empirical data:
- Chi-Square Goodness-of-Fit: $chi^2(50) = 194.26, p < .001$
- Comparative Fit Index (CFI): $ ext{CFI} = .899$
- Root Mean Square Error of Approximation (RMSEA): $ ext{RMSEA} = .078$ (95% CI $[.067; .090]$)
- Standardized Root Mean Square Residual (SRMR): $ ext{SRMR} = .054$
The convergence of acceptable RMSEA and SRMR values confirms the structural validity of the model, supporting the aggregation of the six individual subtest parameters into a unified academic aptitude composite.
Instrument / Measurement Tool
- Official Instrument Name: Berliner Studierfähigkeitstest—Psychologie (BSF-P)
- English Name: Berlin Aptitude Test for Psychology
- Acronym: BSF-P
- Measurement Construct: Psychology Academic Aptitude (Domain-specific scholastic and cognitive ability)
- Test Type: Standardized Cognitive Test Battery
- Administration Mode: Standardized computer-based / electronic testing or proctored paper-and-pencil administration
- Target Population: Upper secondary school students (11th and 12th grades / Abitur candidates) and university applicants applying to undergraduate psychology programs
- Age Spectrum: Adolescents and adults (13–17 years, 18–29 years, 30–39 years, 40–64 years)
- Subtest Composition: Six distinct subtests:
- Reading comprehension (Leseverständnis)
- English language skills (Englischkenntnisse)
- Math skills (Mathematische Fähigkeiten)
- Verbal reasoning (Verbales Schließen)
- Numerical reasoning (Numerisches Schließen)
- Figural reasoning (Figurales Schließen)
- Authentic Response Scale & Scoring Procedure: The BSF-P test score is derived from the average of the z-standardized personal parameters estimated using rapid models for the six sub-tests. Administration can be done electronically or via paper.
- Score Aggregation: Individual subtest responses are first scored for accuracy, converted via IRT models into latent person parameters ($ heta$), transformed into standardized$z$-scores ($M = 0, SD = 1$), and averaged to form the composite BSF-P aptitude metric.
Permissions & Fee and Test Year
- Test Year: 2023 (Operational implementation at Humboldt-Universität zu Berlin initiated in 2021)
- Commercial Status: Non-commercial test battery designed for academic and institutional higher education admissions
- Testing Fee: No licensing fee for qualifying university research and public university admissions
- Permissions & Accessibility: The BSF-P is a secure, high-stakes assessment instrument. To preserve test security, item confidentiality, and the integrity of university admissions, test forms and item pools are not published in the open literature. Qualified academic institutions and researchers must contact the corresponding author to request administrative permissions: Prof. Dr. Matthias Ziegler, Institut für Psychologie, Psychologische Diagnostik, Humboldt-Universität zu Berlin, Unter den Linden 6, 10099 Berlin, Germany ([email protected]).
References
- Bühner, M. (2021). Einführung in die Test- und Fragebogenkonstruktion [Introduction to test and questionnaire construction] (7th ed.). Pearson Deutschland.
- Danner, D., Rammstedt, B., Bluemke, M., & Lechner, C. (2016). Die deutsche Version des Big Five Inventory 2 (BFI-2). ZPID (Leibniz Institute for Psychology Information). https://doi.org/10.23668/psycharchives.940
- Formazin, M., Hell, B., & Ziegler, M. (2011). Studieneignungstests: Konzepte, Konstrukte und Forschungsbefunde [Study aptitude tests: Concepts, constructs and research findings]. Hogrefe.
- Horstmann, K. T., Biesok, A., Witte, K., Godmann, H. R., Fliedner, K., Wilm, L., Doran, L., & Ziegler, M. (2023). Berliner Studierfähigkeitstest—Psychologie (BSF-P) [The Berlin Aptitude Test for Psychology—BSF-P]. Psychologische Rundschau, 74(4), 221–238. https://doi.org/10.1026/0033-3042/a000628
- Jäger, A. O. (1982). Merkmale von Intelligenztests und das Berliner Intelligenzstrukturmodell (BIS). Zeitschrift für Differentielle und Diagnostische Psychologie, 3(4), 273–285.
- Jäger, A. O. (1984). Intelligenzstrukturforschung: Konkurrierende Modelle, neue Entwicklungen, Perspektiven. Psychologische Rundschau, 35(1), 21–34.
- Roemer, C., Ziegler, M., & Bühner, M. (2021). RIAS—Revised Inventory of Academic Interests. Hogrefe.
- Rounds, J., Su, R., Rivkin, D., & Armstrong, P. I. (2010). The academic interests inventory: A new measure of interests for educational and vocational guidance. Journal of Vocational Behavior, 77(3), 361–372. https://doi.org/10.1016/j.jvb.2010.05.006
- Schneider, W. J., & McGrew, K. S. (2018). The Cattell-Horn-Carroll theory of cognitive abilities. In D. P. Flanagan & E. M. McDonough (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (4th ed., pp. 73–163). Guilford Press.
- Watrin, L., Maag, A., & Ziegler, M. (2022). STAV—Studierfähigkeitstest für die Aufnahme ins Bachelorstudium Psychologie [Study aptitude test for admission to the Bachelor of Psychology program]. Hogrefe.
Items of the Scale
Proprietary Notice: The official items, test forms, stimuli, and scoring keys of the Berliner Studierfähigkeitstest—Psychologie (BSF-P) are proprietary and copyrighted. Because the BSF-P is deployed in high-stakes university admissions selection at Humboldt-Universität zu Berlin and collaborating institutions, complete operational item banks are strictly restricted to maintain test security and prevent coaching distortions. In strict adherence to testing standards and ethical testing principles, test items are not reproduced in the open public domain.
Response Scale & Administration Format:
The BSF-P test score is derived from the average of the z-standardized personal parameters estimated using rapid models for the six sub-tests. Administration can be done electronically or via paper.
The battery consists of six core subtests designed to measure distinct facets of academic study aptitude:
1. Reading comprehension (Leseverständnis)
Evaluates the applicant’s ability to comprehend, analyze, and synthesize complex, multi-paragraph German academic and scientific texts relevant to foundational psychology.
2. English language skills (Englischkenntnisse)
Evaluates receptive English competency, syntax parsing, and specialized scientific vocabulary comprehension required to read international psychological literature.
3. Math skills (Mathematische Fähigkeiten)
Evaluates computational proficiency, algebraic problem solving, fraction and percentage calculations, and probability concepts necessary for undergraduate research methods and statistics.
4. Verbal reasoning (Verbales Schließen)
Evaluates crystallized verbal deduction, relational reasoning, logical syllogisms, and semantic analogies within structured linguistic premises.
5. Numerical reasoning (Numerisches Schließen)
Evaluates fluid quantitative induction, rule identification in numerical series, and algorithmic deduction across structured numeric patterns.
6. Figural reasoning (Figurales Schließen)
Evaluates non-verbal fluid intelligence and spatial-inductive problem-solving through abstract geometric matrix configurations and transformation rules.
To obtain authorized testing materials, norm tables, or administrative clearance for institutional research, institutional representatives must contact the primary author team at Humboldt-Universität zu Berlin (Institut für Psychologie, Psychologische Diagnostik).