NR-542 · Week 2 of 8 · Data sources and provenance

NR-542 Week 2 Data Sources and Provenance: How to Write It

The short answer

NR-542 Week 2 asks where the records came from. Provenance is the stage's whole subject: which system produced the extract, what one row represents, who typed each value and under what pressure, what the extract date does to the contents, and which fields were generated by a machine rather than by a person. A data set described without its origins cannot be trusted, and this is the week you learn to write the paragraph that establishes trust. Your section may print this as NR 542 or NR542; it is the same course. Chamberlain publishes no syllabi outside Canvas. The placement here is our teaching judgment from the course's catalog arc; your section's rubric decides what your week actually asks.

NR-542 Week 2 grading scale at Chamberlain, the criterion levels this assessment is scored on, from Chamberlain Tutors
How Chamberlain grades NR-542 Week 2, visualized by Chamberlain Tutors.

What NR-542 Week 2 asks for

Follow one column backwards and the stage explains itself. A staffing report carries a field called acuity, and the number in it drives a conversation about whether a unit was appropriately covered. Trace it back through the extract, through the report definition, through the source table, and it terminates in a dropdown that a charge nurse selects once at the start of a shift, for the unit as a whole, in the ten minutes when she is also assigning patients and hearing about an admission. The value is not wrong. It is simply a rapid categorical judgment made under time pressure by one person, and everything downstream of it inherits that origin, including a committee's confidence in a figure with two decimal places.

That backwards walk is what a provenance section does in writing. Deliverables at this stage typically ask you to identify data sources relevant to a nursing question, describe how the data are generated and captured, and assess what each source is fit for. Some sections add a comparison across sources, such as clinical documentation against administrative or billing data, or device-generated values against manually entered ones.

The distinction worth building the paper around is how a value came into existence. Some values are entered by a clinician as part of care. Some are entered by a clerk for a registration or billing purpose. Some are generated by a device and stored automatically. Some are derived by the system from other fields, which means they carry every weakness of their inputs and none of the visible warning. Those four origins have different error profiles, different completeness, and different meanings, and a paper that sorts its fields into them is doing something most student work never attempts.

One more habit belongs here. Extract dates matter more than students expect, because clinical records keep changing after the moment they are pulled. A discharge disposition captured on day two of an admission is provisional; a diagnosis code assigned before coding review is not final. Saying when the extract was taken relative to the events it describes is often the single most useful sentence in a provenance section.

The NR-542 Week 2 method, step by step

Six moves that establish where a data set came from and what it can be trusted for.

  1. Name the source system and the extract in one sentence each

    What produced the records, what pulled them, when, and covering what date range. Two sentences at the top of the section give a reader the frame for everything that follows.

  2. State the unit of the row and prove it with an example

    One row per encounter, per patient, per event, per administration, per day. Then describe one actual row in words so the definition is demonstrated rather than declared.

  3. Sort every field by how it was created

    Clinician-entered, clerk-entered, device-generated or system-derived. Four short groups, each with a sentence on what typically goes wrong in that group, converts a field list into an analysis.

  4. Describe the capture moment for the fields that matter most

    Who was doing what when the value was recorded, and what else was competing for their attention. This is where nursing knowledge outperforms technical knowledge, and it is the paragraph an analyst could not write.

  5. Say what the source was built for

    Billing data exist to support reimbursement, registration data to identify and admit, clinical documentation to communicate care and meet regulatory requirements. Fitness for your question is judged against original purpose, not against how convenient the file is.

  6. Close with what the set can and cannot answer

    Two lists, short and explicit. Naming the questions this source cannot support is what makes the questions it can support credible, and it sets up the quality work later in the session.

A layout and word budget that establishes trust in a source

Our frame for a data source and provenance paper, sized for roughly 1,100 to 1,400 words. It is our own outline rather than anything the university publishes, and your week's guide outranks it wherever the two disagree.

SectionWhat belongs in itWord target
Question and candidate sourcesThe question you need data for and the two or three sources that could plausibly answer it.140 to 170
Source and extractProducing system, extract mechanism, extract date, date range covered and record count.170 to 200
Unit of the rowWhat one row represents, demonstrated by describing an actual row in prose.120 to 150
Fields by originClinician, clerk, device and derived groups, each with its characteristic failure described.300 to 350
Capture conditionsThe circumstances under which the two or three critical fields are entered, and what that implies.200 to 240
Fitness verdictWhat this source can answer, what it cannot, and which alternative source covers the gap.180 to 220

Evidence craft for provenance writing

Support the claim that a source has known limitations. The behavior of administrative and billing data as a proxy for clinical reality is a studied question with published literature behind it. One citation turns your assertion about coding into a supported premise, and this is the citation most provenance papers are missing.

Describe capture conditions concretely and without blame. A value entered once per shift by a charge nurse during assignment is a description. A value entered carelessly is a judgment you cannot support and do not need. The neutral version is more persuasive and more professional.

Give the extract a date and say what was still changing. Codes assigned before review, dispositions recorded before discharge and orders active at the time of the pull all behave differently from their final values. Naming this is a small paragraph with a large effect on credibility.

Cite any public data set properly and completely. If you are working with a published set, name the issuing body, the release, the reference period and where the documentation lives. Public data sets ship with data dictionaries, and using the dictionary's own field names and definitions is both accurate and easy to verify.

Five mistakes that cost points in this week's territory

  • Naming a file and calling it a source. A filename tells the reader nothing about what produced the records or what a row means.
  • Field lists with no origins. A column inventory is transcription; sorting columns by how they were created is analysis.
  • No extract date. Without it, nobody can tell which values were provisional when the data were pulled.
  • Judging a source against your question only. Fitness has to be assessed against what the source was built to do, or the critique is unfair and unpersuasive.
  • Blaming the people who enter data. Capture conditions explain error patterns; character does not, and the blaming version is both weaker and less accurate.

Before you submit

  • Producing system, extract mechanism, extract date and date range all appear
  • The unit of the row is demonstrated with a described example
  • Every field is assigned to a creation origin
  • Capture conditions are described for the fields the argument depends on
  • The source's original purpose is named before its fitness is judged
  • The section closes with explicit can and cannot lists

Writing the NR-542 provenance section?

Send the data set or scenario and the rubric from Canvas. A premium original draft comes back in 24 to 48 hours with every field sorted by origin and capture conditions described where they matter, and revisions run until the grade lands.

Questions students ask about this stage

How do I describe provenance for a data set I did not build and cannot inspect?
Say what you know, say how you know it, and mark the rest as unknown rather than guessing. A provenance section that states plainly that the extract mechanism is not documented in the materials provided, and that the unit of the row was inferred from the presence of one date per line and repeated patient identifiers, is doing exactly the right thing. Inference is legitimate when it is labeled as inference and supported by what you can see in the file. What loses points is a confident description of a lineage you invented to fill the section. As a bonus, the unknowns you list become the questions you would ask a data steward, which is often material your later sections can use directly.
Which is better for a nursing question, clinical documentation or administrative data?
Neither, in the abstract, which is why the fitness verdict has to be written against a specific question. Administrative and billing data are complete for the encounters they cover, consistently coded, easy to aggregate across settings and structurally blind to anything nobody bills for. Clinical documentation is richer, closer to the patient, and far more variable in completeness and phrasing because it is generated during care rather than for reporting. For a question about volumes, lengths of stay or diagnoses across a system, the administrative set usually wins. For a question about what a nurse assessed, when, and what happened next, only the clinical record can answer, and you will pay for that in messiness. Say which you chose, say what you gave up, and the verdict section writes itself.
Do device-generated values need a provenance paragraph too?
Yes, and students skip them because automatic capture feels self-evidently reliable. It is reliable about different things. A monitor stores what it measured, which is not always what was true of the patient: a saturation value recorded while a probe was off a finger is an accurate record of a bad reading. Automated feeds also have their own gaps, such as intervals when a device was disconnected during transport, and their own volume problem, since a value every minute for six hours is not six hours of clinical assessment. Write one paragraph naming what the device measures, the sampling interval, what happens to the record when the device is disconnected, and whether a human validated the value before it entered the record. That paragraph is often the most sophisticated thing in the paper.

Keep going

Online now