🗂

two views over the same catalog — pick a method, then a project

Pick a method

Browse by family

Project requirements: the deliverable ladder

The project is worth 30% of your grade in the 4-credit version of the course, and completing it is what earns the six sigma black belt certification. You build it across the term in ten steps. Each rung is graded on its own and feeds the next; together they become the final report and the poster. Every step is submitted on Canvas — see there for the current due dates.

New this year: three pairs of 2025 steps have been consolidated into single submissions, marked combined for 2026 below. Ten rungs instead of thirteen — same content, fewer separate deadlines.

  1. Team & topic

    Form a team of 3–5 and commit to one process you can actually observe or get data about. Pick from the catalog above or bring your own — the constraint is access, not ambition.

  2. Your dataset

    Deliver the actual data — experimental, observed, simulated, or found — with a codebook that says what every column means, what its units are, and where each number came from. A plan to find a dataset later does not count.

  3. Project charter

    One page stating the problem, the scope, the metric you will move, and the target. If your team disagrees about the project, the charter is where it shows up.

  4. VOC tree & SIPOC diagram combined for 2026

    Two halves of the same framing move, now one submission. The voice-of-customer tree turns what the customer says into a measurable requirement; the SIPOC bounds the process around it — suppliers, inputs, process, outputs, customers — so you know what is inside your project and what is somebody else's problem.

    Was two separate submissions in 2025.
  5. Process map & literature review combined for 2026

    Open the middle of the SIPOC into the real sequence of steps, decisions, and handoffs — including the rework loops nobody puts on the official diagram — and, in the same submission, establish what is already known about this failure mode or process. Cite what you use; the point is to avoid re-deriving a published result.

    Was two separate submissions in 2025.
  6. Research design

    Say exactly how you will get evidence: what you measure, how often, under what conditions, and what result would change your mind about the cause. Any experiment you intend to run on people gets vetted by an instructor before you run it.

  7. Preliminary results & financial impacts combined for 2026

    First pass at the analysis — the charts, the fitted models, the intervals — reported together with what the improvement is worth in money: cost of poor quality now, cost after, and what the fix would take. Early enough that a broken measurement system is still fixable.

    Was two separate submissions in 2025.
  8. Rough draft

    The whole report assembled end to end while there is still time to fix it. Complete enough to get real feedback on — gaps named as gaps, not quietly left out.

  9. Final report

    The whole arc in one document — define, measure, analyze, improve, control — ending in a recommendation somebody could actually act on. This is the deliverable the black belt certification is judged against, so it is expected to be thorough and complete.

  10. Poster presentation

    The report compressed to what a stranger can absorb standing up: the problem, the evidence, the fix, and the number that proves it worked. Presented live to the class.

Read this before you choose

Two readings: what the project actually asks for, and how to design one that will work. Both are long — expand what you need.

📄 Read: The Project Prompt what the project asks for

For your final project, you will collaborate with 3–5 other students, applying the analytical tools learned in class to real world problems. Completion of the project is required for six sigma black belt certification through this course.

Goals

  1. A relevant question: identify a six sigma black belt design or research project related to systems reliability and/or quality.
  2. Background research: conduct background research and a literature review to define and justify your project.
  3. Design methods: construct an original approach or solution, and apply the reliability and/or six sigma quality tools appropriate to the topic.
  4. Evaluate results: provide a cogent write-up of your work.

Tools and approaches

Use at least one of the six sigma quality control and/or systems reliability models covered in class (below), or get prior approval from an instructor. Define and narrow your topic early so that it can (1) feasibly produce meaningful results in one semester, and (2) name a concrete concept, product, service, or system. Some topics are not covered until later in the term, so talk a topic through with an instructor rather than waiting.

  • The six sigma approach (define, measure, analyze, improve, control)
  • Failure modes and effects analysis (also a six sigma tool)
  • Fault trees and event trees under uncertainty (also a six sigma tool)
  • Component and system reliability
  • Physical acceleration models
  • Statistical process control charts, to detect when performance is deteriorating and to identify root causes
  • Optimization of system design for quality and reliability (e.g. design of experiments; response surface methods)
  • Analysis of variance (ANOVA)
  • Regression analysis

Deliverables

All of them are submitted on Canvas.

Resources

Example abstracts, example papers, and complete reports and posters from past teams are posted on Canvas for enrolled students — see the resources below. To meet the black belt certification requirement, we expect a thorough and complete study.

🧭 Advice on designing your project the four project types

Feeling the aaah of designing your project and looking for some advice? Good news.

Your team & topic is the starting point. The more defined it is, the less of this material you will have to redo later.

Rule of thumb: a great proposal either includes a dataset or a feasible plan to build one. A plan to find a dataset later does not count.

Successful projects generally fall into four types. The first two are particularly recommended.

  • Experimental data in the field
  • Observed data in the field
  • Simulated data using parameters from past studies
  • Observed data from other sources

1 · Experimental data in the field

Summary

You directly randomly assign treatments and record the resulting outcomes. This option can be genuinely fun.

You want easily manipulable treatments that can be tested many times. Think about what products you have access to, what reliability issues you might want to measure, and what treatments you could apply. Then use statistical models, t-tests, and so on to test your hypotheses.

When should I use this?

Experimental data is what you want when you are estimating the value added by a particular choice. Should I use topping A or B on my donut? Which one improves flavor, or willingness to pay, by more?

Baking, queuing, sensor position, short behavioral surveys, willingness-to-pay surveys, shampoo — all easily implemented and cheap. Even "simple" product interventions make very good six sigma projects, because the technique spans many sectors of the economy. You do not have to build a spaceship. What matters is demonstrating the method on a commercially viable product or process.

What analysis does this support?

Experimental data is the gold standard for causal inference in business planning, engineering, healthcare, and science generally. Used with t-tests, ANOVA, factorial design, regression, and response surface methodology — though it works with nearly any tool, so bring other ideas.

Examples

Building your first experiment takes some creativity. Here are starter ideas that a team could implement quickly. (Small bonus points are available as an incentive if you need classmates to participate in your research.)

  • Product problem rating — clothing: how badly does clothing brand X bleed when washed, compared with brand Y?
  • Willingness to pay for product X: your group has a consumer product of interest and wants to know which potential intervention consumers would pay more for.
    • Write a short survey that randomly shows the respondent product version 1 or version 2, then asks how much they would pay for it.
    • Randomly assign respondents to a version and compare results. Ideally, try many combinations.
    • Well suited to consumer-relevant products that are otherwise technically difficult to build real-time treatments for.
  • Process design — food science: six sigma is extremely well suited to making better, cheaper, tastier food.
    • Take a simple recipe and make multiple batches, varying ingredient amounts — bread with a little more flour, or a little less water.
    • Well suited to factorial experiments or response surface methodology, both excellent design-optimization techniques.
    • The same approach applies to chemical experiments. (Instructors volunteer as test subjects.)
  • Product/process outcome — the "no-poo" method: a trend of not shampooing one's hair, which for some people apparently yields an aesthetic benefit. If your loved ones love you very much, recruit a few participants, randomly assign them to a shampoo or no-shampoo regimen over several days, and have them report daily satisfaction. The same philosophy applies to any other commercial product.
  • Product problem rating — ringtones: how annoying is that ringtone, really? A quick behavioral experiment, easily run at scale.
    • Pick several ringtones; randomly assign people to a ringtone and a volume level, and ask them to rate annoyingness.
    • Same philosophy applies to any other commercial product.
  • Factorial experiment — toothbrushes 🪥: buy 12 toothbrushes (3 each of 4 brands) and 12 tubes of toothpaste (3 each of 4 types). Randomly assign a toothbrush and a toothpaste to each group mate, then measure attributes of the brushes and of the outcome after a fixed number of uses.
Class projects generally do not require institutional review board approval. The main rule of thumb is that any experiment you run should be vetted by an instructor before you run it.

Example data format

It often helps to collect data with an online form, so the underlying sheet is shared with your collaborators and fills up in real time.

UnitTreated?OutcomeTrait A
A130 bpm0
A040 bpm1
B050 bpm1
B140 bpm0
more rows hereadd more columns as needed

2 · Observed data in the field

Summary

You collect data in the field, tracking one or more outcomes of interest plus one or more potential treatments or covariates. Unlike experimental collection, the treatment or independent variable is not randomly assigned. That makes statistical control variables essential when you model the data.

When should I use this?

Use it when you cannot randomly assign the treatment — policy, urban planning, observed consumer behavior without intervention. Past teams have, for example, gone into the field and tracked metrics on local transit buses.

Considerations

You might choose this kind of collection if…

  • You are interested in an outcome you can observe but cannot easily or ethically intervene on. For example: which shops see more foot traffic?
  • You want to control the conditions under which data is collected — a random sample of buses, on a random sample of days, at a random sample of locations.
  • You want to study a population with no good existing data that you believe you can reach easily.

Field study ideas

  • Parking: available parking downtown, at the farmers market, or wherever you like is often hard to find. What predicts availability? Randomly select sets of parking spots within an area of interest, randomly select times to check them, and record whether the spot is free, what kind of car is in it, and — if you can stay an hour or more — how long that car stays. If that is too granular, count total usage per street across several days and evaluate the traits of those days.
  • Transit service: the county bus service has visible, trackable metrics. How many people are denied entry because the bus is full? How many are aboard at a given stop? Minutes delayed? Randomly select routes at random times, ride for an hour, and discreetly record non-invasive counts.
  • Water fountains: many campus water fountains are automated and display the total bottles or ounces filled since installation — essentially their lifespan of use to date. Randomly select fountains and stress-test them: how many seconds under the sensor before it detects the bottle? How many tries to fill a bottle to a consistent level?

3 · Simulated data using parameters from past studies

Summary

You use reliability functions, block diagrams, fault trees, and similar tools to approximate a system's reliability over time.

When should I use this?

Simulation is helpful for procedures that are difficult to observe or to test experimentally — expensive products like spacecraft, or sensitive processes like surgeries and medical devices.

What analysis does this support?

You simulate how much the system's reliability, or chance of failure, changes if the mean time to failure of component A shifts by some amount versus component B shifting by another. Which component should we spend money on to fix and improve?

What data must I collect?

Estimates of key parameters for each component in your system — usually the mean time to failure of components A, B, C, and so on.

How do I collect it?

Normally you search past studies and statistics about your device of interest, and find or derive the parameters you need.

Search with Google Scholar. Look for highly cited works published in the last five to ten years, and use the filter to narrow to review articles — a type of article that reviews dozens of prior papers. Some review articles are meta-analyses; they often report dozens of statistics from different studies in a single table and compare them, which is a fast way to assemble your own table of parameters. Here is an example of one such review — its parameters of interest are the sensitivity and specificity of continuous glucose monitors, and it lays out a great deal of study metadata in one table. Many of your projects will want failure rates instead, but the shape of the table is the same.

Tips for data collection

  • Collect any parameter or statistic you can find.
  • The parameters can be of different types — grab any you find. We can often derive one value from another later. For example, any of: mean time to failure; probability of failure at a specific time t; failure rate in failures per hour or per million hours.
  • Always record units — per hour, per thousand hours, per day, number of failed units, probability. Without units the number means nothing.
  • Collect sample size and timespan wherever possible.
    • If a source reports a mean time to failure, that mean came from n observed failures over t total hours. Capture that if you can.
    • If a source reports a probability of failure, you will eventually convert it to a mean time to failure or a failure rate — which needs a sample size n and a timespan t.
  • Collect standard errors and/or standard deviations with sample sizes wherever possible; they are what make an uncertainty analysis possible later. Sometimes they are simply not reported.
  • Not every component will have been researched before. In those cases we either (a) assume it never fails, (b) assume it always fails, or (c) derive a reasonable assumed value and justify the choice. That can be handled later.
  • Always download a copy of the document and its citation information.
  • Always cite where the estimate came from.
  • If you have access to raw data you could estimate the value yourself, but often you will not.

Your dataset will look roughly like…

Component / failure modeParameterEstimateStandard errorSample size (N)Source
Selfie stick armMean time to failure3000 hrs150 hrs200Source 1
Camera imagingFailure rate5 per million1 per million50Source 2
Clasp mechanismProbability of failure at 1000 hrs5%1.2%1200Source 3
etc. — expect at least 5 rows, ideally more

4 · Observed data from other sources

Summary

Statistical analysis of pre-collected data, tracking metrics for multiple products over time or space under varying conditions. Any features, treatments, or interventions were observed out in the world rather than randomly assigned — which means it is not appropriate to draw causal conclusions directly from observed data.

When should I use this?

Observed data is useful when a topic has been studied extensively in the public domain, when there is little or no individual privacy issue or commercial risk in sharing it, and when collection systems are already well developed. Statistical models such as regression are used to study it. Good for initial exploratory analysis, but experimental designs are preferred where possible. Usually needs at least 100 observations, and more will sharpen your estimates.

A statistical model of this data can still identify which features improve quality in a specific context. It generally requires quality metrics observed over time for multiple groups — EV battery lifetimes over time by vehicle, say, or hot spring performance over time.

What data must I collect?

The entire dataset, meeting these requirements:

  • It must be completely clear what a row means and what each column or variable means.
  • Units must be clear for every variable used — otherwise the analysis is uninterpretable.
  • There must be variation in the observed traits, or there is no variation to analyze.
  • Sample size must be large enough. At least 100 rows is a good benchmark; higher is better.
  • The source and quality of the data must be good. If it is not clear who collected it, how, and when, it is not good enough to present to a client — and that is what this project is about.

Tips for data collection

  • Using pre-collected data may look like the easiest option. It often is not.
  • Data quality varies enormously, and the goal is not merely to run a statistical model — it is to produce an empirically valid, useful analysis of a technology for the improvement of quality.
  • Only take this route if you have a high-quality dataset that is meaningfully connected to your project and has a clear, valid, well documented source.

Your dataset will look roughly like…

Unit IDTimestepOutcomeTrait AMore traits here
A132 bpm0
A235 bpm0
B167 bpm1
B268 bpm1
more rows here

Examples from past projects

Both of these are past students' own work, so they live on Canvas rather than on the public site. Log in with your Cornell credentials and open the course modules.

📝 Example project abstracts

Short abstracts from previous teams, showing how a good project states its question, its method, and its result in one paragraph. Posted in the course modules for enrolled students.

Open on Canvas →

🏅 Past black belt projects & posters

Complete reports and posters from previous teams — the standard a certification-level study is held to. Posted in the course modules for enrolled students.

Open on Canvas →