Skip to content
tezvyn:

How would you architect a centralized research repository for heterogeneous data types?

Source: nngroup.comMediumHow cards are made

How would you architect a centralized research repository for heterogeneous data types?

Tests designing systems that unify unstructured UX artifacts into a queryable graph. Cover: ingestion with transcription, a metadata schema linking insights to evidence, and faceted cross-type search. Red flag: a flat file dump without structured tagging.

What's really being asked

This question evaluates your ability to design a socio-technical system that transforms raw heterogeneous UX artifacts into a structured discoverable knowledge base. The interviewer cares about your understanding of data pipelines, information architecture, and search infrastructure in a research operations context, not just your familiarity with UX methods or report writing.

The full answer

A strong response walks through four layers in order. First, ingestion: describe automated transcription for video, parsing and normalization for survey responses, and entity extraction so every artifact becomes machine-readable. Second, storage: propose a unified metadata schema where insights, transcripts, clips, and raw notes link back to studies and participants via persistent identifiers. Third, indexing: explain a search index that supports full-text search across transcripts and tags, plus faceted navigation by study date, method, participant segment, or theme. Fourth, cross-referencing: show how relationships are explicitly modeled, such as linking a survey quote to a video timestamp and the insight it supports, so users can trace evidence from claim to primary source.

The mistakes people make

Red flags include proposing a simple folder hierarchy in a generic drive like SharePoint with no structured metadata, which replicates the silo problem. Another mistake is optimizing only for video search while treating surveys as static PDFs, missing the need for unified querying across qualitative and quantitative data. Suggesting manual tagging without automation also signals a design that will not scale past a handful of studies.

What usually comes next

Interviewers often push on versioning when insights evolve, access control for sensitive participant data, or how to prevent the repository from becoming an uncurated data swamp. They may also ask how you would measure adoption, trust, or return on investment for the ResearchOps team.

A concrete example

Imagine a repository receiving a usability test. The ingestion pipeline transcribes the session video and extracts timestamped utterances. A researcher tags a segment as checkout friction. The system links that tag to a related survey response where twelve participants rated checkout ease as low. When a product manager later searches checkout friction, they see the insight, the video clip, and the quantitative survey trend, all cross-referenced under a shared theme taxonomy that updates across studies.

Interview question

Which architectural approach best distinguishes a queryable research repository from a siloed flat file dump when unifying heterogeneous UX artifacts?

  • a.Organizing all files in a single cloud drive with a strict folder hierarchy by study date and method
  • b.Designing a unified metadata schema that links insights, transcripts, and source data through persistent identifiersCorrect
  • c.Requiring researchers to manually tag every artifact with standardized keywords before ingestion
  • d.Processing videos with automated transcription while storing survey responses as static PDFs
Why?

A unified metadata schema with persistent identifiers forms the storage layer that explicitly links insights to heterogeneous evidence, enabling faceted cross-type search and traceability. Option A seems organized but is essentially a flat file dump without structured relationships, which is a key red flag.

Just read this? Test yourself on what you have been reading.

Read the original → nngroup.com

You just looked this up. Could you explain it out loud?

That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.

The iPhone app is on the way

We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.

Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.

Get it on Google PlayiPhone app coming soon

We are hiring for this. Open roles that interview on ux research — each one lists the topics its interview covers.

See open roles