Project S: ← Home

Where it all came from

Summer 2026. A heatwave the likes of which Warsaw hadn't seen in years. During the day, work is simply impossible, so the evenings are given over to entertainment instead — that is, to tasks nobody can say for certain will ever benefit anyone. Especially in research, where data quality has for some time now seemed to matter less and less.

The direct impulse behind the program was a sight no methodologist should have to witness without prior psychological preparation: students entering survey data straight into the cells of a popular office program, or their marginally cleverer — though still not clever enough — counterparts, copying it out by hand into a web form instead. The second solution, admittedly a touch better, remains almost the very same error generator, just in new packaging, with the added bonus of a significantly longer working time. Time that could far more sensibly have been spent on proper data validation instead of mindless re-typing.

Someone might rightly point out here that data-entry programs have existed for a long time and have their own, at times venerable, history. True enough. Among domestic solutions worth mentioning is the legendary WD (not to be confused with the bicycle-chain lubricant, though on data chains it worked rather like grease on a bicycle), prepared by Wojciech Niepokojczycki for LUTAY H. C. Zbigniew Sawiński, known for his enormous contribution to the development of survey data processing (and not only that!) — or the later Windows-based DGS by YAC Software, which drew on the legacy of WD. Besides various conveniences, it kept backward compatibility with the WD language syntax, which placed it among the direct successors and continuators of that proud tradition. Both programs were, in their time, used by what was probably the majority of research agencies in Poland. Some of their limitations — such as the simplified, single-condition syntax and layered funnel filters limiting the possibility of advanced routing — were eliminated in later development versions of Puncher, produced by URUSoft, which was used, for instance, to load every edition of the Polish General Social Survey from 2005 onward.

A separate category of data-entry programs consists of tools that are genuinely powerful, but so complex that mastering their own language and structure takes more time than their use could ever save. In simple projects, this makes them comparable to a pneumatic hammer used to drive in a thumbtack: technically effective, practically disproportionate to the task. Students therefore cope fairly rationally — especially since, today, free data-entry tools aren't exactly abundant on the market, as most of them have vanished, leaving only a fond memory behind. DE is an answer to that gap: a free (for students), simple program for entering data from paper questionnaires, similar to existing solutions but enriched with features designed to improve work efficiency.

At the very same time DE was taking shape, a friend at GESIS (Leibniz-Institute für Sozialwissenschaften), Petra Brien, who works in the Department of Survey Data Curation and was reviewing data from the Polish edition of the ISSP Digital Societies study, reported to me an observation worthy of Holmes himself: two records in the Polish dataset looked almost identical, yet their degree of similarity narrowly failed to reach the threshold at which the duplicate-detection script developed by GESIS DSDC would have flagged them as such. The GESIS tool detects duplicates with 100% similarity as well as near-duplicates above a 96% threshold (excluding demographic variables). GESIS has, incidentally, been dealing with the issue of record duplicates in ISSP since 2014. A broad discussion in the ISSP Methodology Committee, held during the General Assembly in Guadalajara in 2018 and prompted, among other things, by the findings of a report submitted by the GESIS Archive, authored by Insa Bechert, Petra Brien and Marcus Quandt, together with its recommendations, resulted in duplicate analyses being extended to include near-duplicate analysis as well, following the methodology developed by Kuriakose and Robins after its adaptation to the structure of ISSP datasets. At the same time, national PIs were obliged to carry out the relevant tests themselves, which is documented in Technical Reports (formerly the SMQ — Study Monitoring Questionnaire), describing the methodological and operational context of the national studies. The GESIS tools, it's worth noting, still do their job excellently to this day.

A deeper reflection on the problem inevitably led me to the conclusion that the phenomenon of duplicates in datasets can be viewed in at least two ways. The first is threshold-based: a — in essence, arbitrary — cutoff is set, above which two sufficiently similar records are declared duplicates. This assumption is consistent with the logic of an audit, in which a boundary allowing a positive or negative qualification must be defined. The second view, in my opinion more realistic, and more important from the standpoint of data quality control, grounded in the Kuriakose and Robins methodology's emphasis on degree of similarity, holds that this degree is at the same time a measure of the informational redundancy those records — and data more broadly — carry with them. The more similar two records are, the less new information the second one contributes relative to the first. If that's the case, the degree of similarity is better expressed on a continuous scale, rather than reduced to a single cutoff above or below which similarity is or isn't accepted.

Turning this thought toward Petra's observation proved, for me, more costly than I had expected — the rest of the evening went to pondering how one might solve the problem in a way that, beyond simply setting an acceptance threshold in line with ISSP's methodological requirements, would let the tool examine the entire dataset, and any part of it, as broadly and flexibly as possible — including in terms of record similarity.

That flexibility would open up at least two ways of assessing data quality methodologically. First, it would make it possible to identify the parts of a questionnaire that carry the weakest informational value. Second, it would make it possible to detect near-duplicate records, while also flexibly treating the structure of the dataset itself — something already achieved, to some extent, by the tool developed by GESIS and built on an external Stata-based module authored by Kuriakose himself (2015).

From this small reflection grew an ambition far greater than merely tracking down duplicates. If suspicious pairs of records can be diagnosed, the response patterns themselves can also be studied — especially within batteries of questions measured on identical scales. This is, of course, nothing new, though it is rarely used in the context of assessing the informational value of data. Along the way, further diagnostic parameters were added to the set, which turned out useful not only for catching anomalies but also for ordinary data analysis.

The idea eventually grew into a project of its own: Audit 2.0 — a module capable of analysing data from other sources too, not necessarily entered directly through the program, though that remains its primary calling. And so, out of DE, DEVS was born: the same tool, only enriched with diagnostic and control functions for assessing data quality — including the less obvious kind, tied to informational value. This thread will be continued, and substantially expanded, in future tools, in particular the planned Project S: QA. And that is also how Project S: itself was born — the umbrella under which both tools were created, with at least a few more planned, all built around successive dimensions of data-quality assessment.

The Ancestor's Spirit

A similar thought had visited the author before — not for the first time. Back in 2002, together with Dr. Tomasz Jerzyński, we created Validator 1.0 — a program used in the Polish General Social Survey 2002, whose main task was to carry out advanced logical checks on the dataset and validate the questionnaire itself against its transition and skip rules, as well as advanced checks of variables unrelated to the tool's own logic. This was no simple task, either, since a significant part of those rules was expressed as descriptive instructions for the interviewer, interpreted on the spot. Those were, however, the days when instruments still took paper form — far more flexible than algorithmised CAPI or mobile scripts. Working with them also demanded far better professional preparation from interviewers than is the case today. The rules governing the course of the interview therefore first had to be reconstructed, bearing in mind the logic not only of the interview itself, but also of the respondent, and even of the interviewer. Those were, in their own way, beautiful times, when studies, owing to their apparent algorithmic imperfection, forced a genuine understanding of the interview situation, have sadly faded into obscurity, giving way to soulless scripts and automation — losing, in the process, the very essence of what they were supposed to be about.

Attempts to turn Validator 1.0 into a more universal tool, however, ended in a failure worthy of a book of its own. When adapting it for the 2005 edition of PGSS, it became clear that basing the program on rigid, static code had been a strategic mistake — every subsequent change, even a minor one, in the questionnaire's structure consumed more effort to adapt the program than it could ever pay back in saved work. The same goals could have been achieved faster and more effectively by an entirely different route. The program was thus laid quietly to rest, never to see a successor in the form of another edition.

The idea itself — combining automatic validation with the very act of data entry — was nonetheless a sound one, and it worked splendidly in 2002. Which proves that good ideas can outlive the failure of the tool that first embodied them.

DEVS, together with the Audit 2.0 module, is therefore, in a sense, Validator's grandchild: the same idea, this time built on a completely different philosophy than two decades earlier, and enriched with at least several dozen other features that the authors, building Validator more than twenty years ago, never even thought of.

And what about those two records?

Here the story serves the reader its most ironic twist of all: the cases Petra reported were ultimately removed from the ISSP Digital Societies dataset — but not because of the suspected duplication at all. Further analysis revealed entirely different quality issues.

The troubling, if insufficient, similarity turned out to be an alarm signal — a false lead that nonetheless led to the actual culprit behind the whole affair. The mechanism thus performed flawlessly, just not in the way anyone had expected, which, it seems, makes for the best possible origin story for a tool built to detect anomalies. The study data, including the Polish part, can be downloaded from https://doi.org/10.4232/5.ZA10020.1.0.0 (ISSP Research Group, 2026. International Social Survey Programme: Digital Societies I - ISSP 2024. ZA10020; Version 1.0.0. GESIS, Cologne) — naturally, without the two flawed records from which it all began.

To spare lecturers (and above all the data itself) unnecessary suffering, any student who needs a data-entry tool can obtain a free licence to use the programme free of charge. Simply fill in the form, giving your full name, university, faculty or institute, a university email address, and the purpose for which the program will be used — and wait a few days.

Marcin W. Zieliński
Summer 2026