He Had Done His Thinking
Dr. Junaid Rashid arrived before anyone else.
This was new. In the early sessions, he had arrived exactly on time, with the punctuality of a man with fifteen years of ward discipline behind him. Then he started arriving five minutes early. Then ten. Tonight, I found him already in the Research Room, his notebook open, a diagram drawn in blue pen across two full pages.
It was a timeline.
On the left: a group of people. Some marked with a small tick, exposed. Some marked with a small cross, not exposed. An arrow ran from both groups across the page to the right. At the far end of the arrow, some of the people from both groups had developed a circle he had labelled “disease”. Others had not.
He had drawn a cohort study. Without being told what one looked like.
“Sir, you said the cohort moves forward,” he said. “So I drew it moving forward.”
“How long did this take you?”
He thought about it. “Forty minutes. I kept redrawing the arrows.”
“The arrows are correct,” I said. “Explain it to me.”
The Study That Follows Time
By the time the others arrived (Dr. Sumaira Talib with her usual notebook, Dr. Hammad Ali slightly breathless, Dr. Hassan Raza reading something on his phone until the last possible second), Junaid had already been talking for three minutes. I let him.
“In the cross-sectional study,” he said, pointing at the whiteboard where I had redrawn his diagram in larger form, “you take one photograph. Everyone is measured at the same moment. You cannot see time. In the case-control study, you start at the outcome and look backwards. You are reconstructing the past from memory.”
He paused.
“But in the cohort, you start before the disease happens. You take a group of people who do not have the disease. You identify who is exposed to a risk factor and who is not. Then you follow both groups forward, for weeks, months or years, and you watch who develops the disease and who does not. Because you were there before the disease appeared, you know the exposure came first. The timeline is real. Not reconstructed. Not remembered.”
Sumaira set down her pen. “That is the closest thing to proof,” she said quietly.
“It is the closest an observational study can get,” I said. “You cannot run a randomised trial on smoking. You cannot assign someone to smoke for twenty years and wait. But you can follow smokers and non-smokers and watch what happens. The cohort study is how we know most of what we know about the major risk factors in modern medicine.”
What the Cohort Actually Measures
I turned to the board. “Because you start with people free of the disease and count who develops it, a cohort gives you something the cross-sectional and case-control designs cannot: incidence. Divide the new cases by the number of people you started with, and you have the risk in that group over the follow-up period. A randomised trial also follows people forward, so it gives incidence too. Among observational designs, the cohort is the one that does.”
“And from incidence comes the measure of association. Not the odds ratio this time. The Relative Risk, also called the Risk Ratio, written as RR.”
RR = risk (cumulative incidence) in the exposed group ÷ risk (cumulative incidence) in the unexposed group
Hassan uncrossed his arms, which is his version of leaning forward. “So an RR of 2.5 means…”
“The exposed group is 2.5 times as likely to develop the disease as the unexposed group, over the same follow-up period. An RR of 1 means no difference. Below 1 suggests the exposure is protective. And like every estimate, you report it with its 95% confidence interval.”
“So this is the number the odds ratio was only approximating,” said Junaid.
“When the outcome is rare, yes. We covered that gap in the case-control session. In a cohort you do not need the approximation. You measure the risk directly.”
I added one more line under the formula. “When patients join your cohort at different times and are followed for different lengths of time, you care about when the event happens, not only whether it happens. Then researchers report a hazard ratio, usually from Cox regression. You do not need to calculate it by hand. Know the name when you see it in a paper, and ask a statistician when you need it.”
A Worked Example on the Board
Hammad looked uneasy. “Sir, can you show us with numbers?”
So I drew a 2×2 table, using Junaid’s own patients as an imaginary example. Three hundred hypertensive patients from the medicine OPD, none with a previous crisis. At the start, 150 are non-adherent to their medicines and 150 are adherent. All are followed for one year. The outcome is a hypertensive crisis needing admission.
| Group | Crisis | No crisis | Total | Risk |
|---|---|---|---|---|
| Non-adherent (exposed) | 30 | 120 | 150 | 30 ÷ 150 = 0.20 |
| Adherent (unexposed) | 10 | 140 | 150 | 10 ÷ 150 = 0.067 |
RR = 0.20 ÷ 0.067 = 3.0
“In this made-up example,” I said, “non-adherent patients had three times the risk of a hypertensive crisis over one year compared with adherent patients. Twenty percent against about seven percent. That is a sentence a physician can use in the OPD tomorrow.”
The Framingham Moment
I put the marker down and told them a story.
In 1948, in a small town called Framingham in Massachusetts, the United States Public Health Service enrolled 5,209 adult men and women to study heart disease before it appeared [1,2]. They were examined. Their blood pressure was measured. Their weight, their smoking habits, their cholesterol and their family histories were all recorded. Then they were sent home and asked to come back for repeat examinations every few years.
They kept coming back.
Their children were enrolled in 1971. Their grandchildren in 2002 [2]. The Framingham Heart Study has now run for more than seventy years. From that cohort came much of what we teach about high blood pressure, high cholesterol, smoking, diabetes and obesity as causes of heart disease. In 1961, its investigators described these as “factors of risk”, and the term “risk factor” entered everyday medicine [1,2].
“Before Framingham,” I said, “many doctors thought heart disease was simply part of ageing. Something that happened to you. The cohort study changed that. It gave us the risk factor, and risk factors can be treated.”
The room was quiet in the way it gets when something larger than methodology has entered the conversation. Junaid was looking at his diagram. I could see him placing his forty minutes of blue-pen arrows inside that history.
The Two Kinds of Cohort Study
“Before the limitations,” I said, “one distinction.” I drew a dividing line on the board.
- Prospective cohort: you recruit participants now, before the outcome has occurred, and follow them forward in real time. The strongest form, and the most resource-intensive.
- Retrospective cohort: you identify a group assembled in the past (a patient register from 2018, a hospital database, a birth record) and trace them forward through existing records to see who developed the outcome. You still move forward through time, but on data collected earlier.
Hammad frowned. “But a retrospective cohort and a case-control both use old data. What is the difference?”
“The starting point. In a retrospective cohort, you start with exposure status (who was exposed and who was not) and follow them forward through the records to the outcome. In a case-control, you start with outcome status and look back for exposure. The data may be equally old, but the direction of reasoning is opposite. Reviewers spend their lives catching that mistake. Do not label your study by the age of the files. Label it by where you start.”
The Limitations That Make It Expensive to Love
Sumaira turned a page in her notebook. “Now the difficult part.” She had learned to expect it. Every design has one.
Loss to follow-up. “This is the cohort study’s greatest enemy,” I said. “You enrol five hundred patients today. In year three, one hundred have moved. In year five, eighty more have stopped answering the phone. And the people who leave are almost never a random sample of those who stay. Sicker patients drop out. Patients who moved to another city drop out. Your remaining cohort no longer looks like the one you started with.” I wrote it on the board: attrition bias. “You reduce it with careful follow-up, two phone numbers per patient, reminder calls and appointments matched to clinic days. You cannot remove it completely. You report how many were lost, and you compare the dropouts with those who stayed, so the reader can judge whether they differed.”
Confounding. “You did not choose who was exposed. The patients chose, or life chose for them. Non-adherent patients may be older, poorer, or smoke more. Maybe it is the smoking, not the missed tablets, that causes the crisis. A cohort shows you that the exposure came first. It does not, on its own, prove the exposure caused the outcome. So you measure the likely confounders at the start, adjust for them in the analysis, and still say honestly that unmeasured confounding may remain.”
Time and cost. “A prospective cohort of a disease with a ten-year latency needs ten years of funding, staff and data management. In Pakistan, where research funding is already thin, choose this design when the question justifies the investment and your institution can sustain it.”
Rare diseases. “A cohort of two thousand people followed for five years will not produce enough cases of a disease that affects one in ten thousand. For rare outcomes, the case-control design remains the efficient choice.”
The Hawthorne effect. “When people know they are being watched, they sometimes behave differently. A patient told to return every month for a blood pressure check may suddenly start taking his tablets more carefully. Your cohort’s behaviour can shift simply because it is being observed.”
The Cohort a Trainee Can Actually Finish
Sumaira looked worried. “Sir, a trainee has a few years at most. Nobody in this room can follow patients for ten years.”
“You do not have to,” I said. “The realistic cohort for a CPSP trainee is a short, hospital-based prospective cohort. Think of Dr. Bushra Fatima’s synopsis. She is on duty in the labour room tonight, but her question fits this design well. Preterm babies whose mothers received antenatal corticosteroids are the exposed group. Preterm babies whose mothers did not are the unexposed group. Both groups are followed from birth until discharge from the nursery, and she records outcomes such as respiratory distress and NICU admission. The follow-up is days, not years. The patients are in her own hospital. Loss to follow-up is small.”
“And because she compares two groups, her synopsis needs a non-directional hypothesis, as we discussed in the objectives session. Follow your supervisor and the current CPSP synopsis guidelines on the rest. But remember: a cohort does not always mean years.”
“One more thing,” I said. “She cannot ethically decide which mothers receive steroids. That is a clinical decision already made by the treating team. She only observes it. The moment a researcher assigns the exposure, it is no longer a cohort. It is an experiment, and that is Hammad’s territory.”
Hammad sat up straight, delighted to have territory.
The Diagram That Connected Three Sessions
I turned back to the whiteboard, where Junaid’s diagram had grown over the hour. Three designs now stood side by side.
| Design | Direction | Main measure | Best for |
|---|---|---|---|
| Cross-sectional | Single snapshot | Prevalence | How common is this? |
| Case-control | Backwards from outcome | Odds ratio | What factors are associated with this outcome? (rare outcomes) |
| Cohort | Forward from exposure | Incidence and relative risk | What will happen to this exposed group, compared with an unexposed one? |
I looked at Junaid. “You started three sessions ago wanting to know how many hypertensive patients in your OPD were non-adherent. Cross-sectional. That question grew into asking whether non-adherence was linked to the crises you admit. Case-control. And now?”
He looked at his diagram. He already knew.
“Now I want to know what happens to my non-adherent hypertensive patients over the next three years, compared with patients who take their medicines regularly. How many in each group have a hypertensive crisis, a stroke, or a hospital admission.”
“You found the most important words yourself,” I said. “Compared with. A cohort needs an unexposed group. If you follow only non-adherent patients, you have nothing to compare their risk with.”
In those sentences was a prospective cohort study: a publishable, clinically useful piece of work that could change how physicians in Pakistan counsel their patients about adherence. One clinical problem. Three designs. Three questions, each deeper than the last.
“Write that down,” I said. “Exactly as you said it.”
He already was.
What Happened When the Session Ended
The others left first. Hammad was animated, telling Hassan about the old surgical register in his unit and whether it could hold a retrospective cohort. Sumaira paused at the door and said, without looking back, “These three sessions belong together. As a set.”
She was right.
Junaid stayed, as he always does. But tonight he did not stay to ask a question. He stayed to finish the diagram, adding the measure of association under each design, and the word “incidence” in a box at the end of the cohort arrow, with a circle around it.
When he finally closed the notebook, he held it for a moment before putting it in his bag. Fifteen years of watching disease. Of treating it and documenting it in files no one would ever analyse. Fifteen years of data that had never been asked the right question.
The notebook was starting to ask.
“Same time next week, Sir?”
“Same time next week,” I said.
Key Takeaways
- A cohort study follows an exposed and an unexposed group, both free of the outcome at the start, forward in time. Without a comparison group, it is not a cohort.
- It shows that the exposure came before the outcome and can study several outcomes from one exposure.
- It gives incidence and the relative risk: RR = risk in exposed ÷ risk in unexposed (for example 0.20 ÷ 0.067 = 3.0), reported with a 95% confidence interval. For time-to-event data, researchers report a hazard ratio from Cox regression.
- Prospective and retrospective cohorts both start from exposure; a case-control study starts from the outcome.
- The main limitations are loss to follow-up, confounding, time, cost and poor efficiency for rare outcomes.
- For a CPSP trainee, a short hospital-based prospective cohort, like Bushra’s antenatal corticosteroid study, is the realistic choice.
References:
- Mahmood SS, Levy D, Vasan RS, Wang TJ. The Framingham Heart Study and the epidemiology of cardiovascular disease: a historical perspective. Lancet. 2014;383(9921):999-1008.
- Tsao CW, Vasan RS. Cohort Profile: The Framingham Heart Study (FHS): overview of milestones in cardiovascular epidemiology. Int J Epidemiol. 2015;44(6):1800-13.
Need personalised help with your synopsis, data analysis, or manuscript? UPMED Medical Consultancy | WhatsApp: 03042397393 | [email protected]
Follow the UPMED Medical Consultancy Channel to stay updated on The Research Clinic: A Doctor’s Journey from Question to Publication series. We will share posts covering all the latest updates and progress. Join the WhatsApp Channel
You can also connect with the writer of this blog post series to share or receive suggestions: Dr. Junaid Rashid (Founder of UPMED) | WhatsApp: 03042397393
List of all the posts in this journey.
