Intellicure Analytics Logo

7 Strategies for a Credible Skin Substitute Efficacy Comparison Using Real-World Data

Oct 6, 2026

CMS’s CY2026 Physician Fee Schedule final rule changed how skin substitutes are paid. Coverage has moved too: the Medicare Administrative Contractors (MACs) withdrew the unified skin substitute local coverage determinations (LCDs) for diabetic foot ulcers (DFUs) and venous leg ulcers (VLUs) days before their January 1, 2026 effective date, and in July 2026 CMS quietly reopened its call for clinical evidence on these products. Because payment and coverage policy are both still in motion, confirm current details at cms.gov. What does not change is the evidentiary burden. Manufacturers defending a product, payers writing policy, and value analysis committees choosing among products all need a skin substitute efficacy comparison that survives audit and peer review.

Most comparisons fail on method, not intent. Products are pooled by code, etiologies are blended, and unadjusted healing rates become rankings. The seven strategies below form a method checklist, not a product ranking, and each states the data limitations plainly: confounding by indication, documentation variability, right-censoring, and imperfect code-to-product mapping.

It’s important to note that developing a Comparative Effectiveness Study with Intellicure Analytics makes all of these steps easy, or simply completed as part of the process. Click here to learn more.

1. Define a comparable product set

Every downstream result inherits the quality of the product definition. Skin substitutes span different classes (for example, placental allografts, xenografts such as bovine or porcine collagen matrices, and synthetic or composite matrices) and different regulatory pathways: 361 human cells, tissues, and cellular and tissue-based products (HCT/Ps) regulated under section 361 of the Public Health Service (PHS) Act, devices cleared via 510(k) or approved via PMA, and biologics under a BLA. CMS now uses these same FDA categories to group skin substitutes for payment, so pathway is more than a regulatory footnote. Confirm current pathway distinctions at fda.gov. Products on different pathways carry different evidence expectations and different mechanisms, so treating them as interchangeable produces a number nobody can interpret.

Illustration: suppose you want to compare a placental allograft with a bovine collagen matrix in DFUs. Pooling them into a generic “skin substitute” arm would hide any difference between them. Separating them by class first, then reporting within-class results, lets a reader see which comparisons are like-for-like and which are exploratory.

The central difficulty is HCPCS crosswalking. Q-codes (and, for some products, newer A-codes) are billing constructs, and a single code can map to more than one product, or a product can change codes over time. Claims-derived and EHR-derived data both inherit this problem.

  1. Build a product dictionary with one row per unique product.
  2. Map each Q-code or A-code to product names and effective date ranges, since mappings change.
  3. Assign class and regulatory pathway to every product, with the source documented.
  4. Set a minimum number of wound episodes per product before it enters any comparison.

The common mistake is pooling products that share a Q-code, or comparing across unlike categories as though they were substitutable. Where a code cannot be resolved to a single product, report it as unresolved rather than guessing.

Measure the percent of records mapped to a unique product and the number of episodes per product. A low mapping rate is itself a finding about what the data can support.

2. Stratify by wound etiology and baseline severity

A pooled healing rate across DFUs, VLUs, and pressure ulcers is not a measure of product performance. These etiologies differ in biology, expected trajectory, and standard of care, and a product’s case mix across them varies widely. A product used mostly in smaller, newer DFUs will look better than one used in large, chronic VLUs regardless of what the product does.

Illustration: suppose Product A shows a higher unadjusted closure rate than Product B. When baseline wound area is balanced between the groups, the gap shrinks to nothing, because Product A was applied disproportionately to smaller wounds. The apparent superiority was a property of the patients selected, not the product.

The way to prevent this is to fix the analytic structure before touching outcomes:

  1. Define baseline covariates: wound area, wound duration, depth, and relevant comorbidities.
  2. Set strata by etiology, and analyze each separately.
  3. Apply matching or weighting within strata.
  4. Check covariate balance with standardized mean differences (SMDs) before opening the outcome analysis.

The mistake to avoid is reporting a single pooled rate across etiologies. A related trap is adjusting for baseline variables that are poorly documented, which trades one bias for another.

Measure covariate balance after weighting. An SMD below 0.1 is a common convention, not a guarantee of adequate control. Also report missingness of each baseline variable, since area and duration are often incompletely recorded in routine care and the handling of that missingness (imputation or complete-case) should be stated.

3. Pre-specify the efficacy endpoint

Endpoint choice is where results get shaped, deliberately or not. If you can choose among complete closure at 8, 12, or 16 weeks, percent area reduction, or time to closure after seeing the data, some definition will favor your preferred product. Writing the definition down first removes that degree of freedom, and reviewers know it. Complete wound closure also remains the FDA’s preferred primary endpoint, with percent area reduction alone generally treated as supporting rather than primary evidence, a point the Wound Care Collaborative Community recently revisited in its proposed addendum to FDA’s chronic wound guidance.

Illustration: run complete closure at 12 weeks and percent area reduction at 4 weeks side by side. If the two agree on direction, the conclusion is more robust. If they disagree, that disagreement is informative and should be reported rather than resolved by picking the better-looking one.

The protocol should specify:

  • The closure definition, including what documentation counts as closed (for example, a clinician-recorded closure versus a measured area of zero).
  • The index date, typically first application.
  • The follow-up window and the rule for episodes that end early.
  • Censoring rules for transfer, death, discharge, or loss to follow-up.

Real-world follow-up is rarely complete, so a simple “healed by week 12” proportion treats unobserved patients as failures or drops them, and both distort results. Use time-to-event methods such as Kaplan-Meier or Cox models, which handle right-censoring explicitly, and consider competing events.

The common mistakes are switching endpoints after seeing results and ignoring loss to follow-up. Measure the proportion of episodes with adequate follow-up and the agreement between primary and secondary endpoints.

4. Use active-comparator real-world designs

Head-to-head randomized trials between skin substitutes are scarce, so many decisions rest on observational data. The credible form of that evidence is a new-user, active-comparator design: patients starting product A are compared with patients starting product B, rather than with untreated patients or prevalent users. This aligns time zero, avoids survivor bias from patients who tolerated earlier applications, and compares people who were all candidates for advanced therapy.

Illustration: a comparative effectiveness study of two cellular and/or tissue-based products (CTPs) in DFUs includes only new users, weights the groups on baseline covariates, and states plainly that unmeasured confounding (such as perfusion, glycemic control, or adherence not captured in the record) may remain.

  1. Define the index date as the first application of the study product.
  2. Exclude patients with prior use of either product in a lookback window.
  3. Fit the propensity model on pre-index covariates only.
  4. Apply propensity or overlap weighting, and verify balance.
  5. Estimate effects with robust confidence intervals.
  6. Run an E-value or similar sensitivity analysis to quantify how strong an unmeasured confounder would need to be.

The mistake is presenting observational results as equivalent to randomized evidence. Confounding by indication is the central threat: clinicians choose products for reasons the dataset may not record. Say so, and word conclusions as associations.

Measure the stability of the effect estimate across weighting methods (for example, inverse probability versus overlap weighting) and report the sensitivity-analysis outputs. An estimate that flips under a reasonable alternative specification should not be headlined.

5. Account for care setting and utilization patterns

Skin substitutes are applied in skilled nursing facilities (SNFs), in home visits, in hospital-based wound centers, and in offices. These settings differ in patient acuity, visit frequency, documentation habits, and the standard of care delivered between applications. They also differ in how products are paid: the 2026 flat rate applies in offices and hospital outpatient departments, while products used during a covered Part A SNF stay or a home health episode fall under bundled payment. A comparison that ignores setting can attribute a setting effect to a product, and pooled rates across settings are not comparable.

Utilization patterns add a second layer of information that outcomes alone miss. Illustration: suppose a product shows a high number of applications per episode and frequent switching to other products. That pattern more plausibly signals non-response and escalation than efficacy, yet raw utilization volume is often read as a sign that clinicians favor the product.

  1. Tag each episode by care setting at the index date.
  2. Compute applications per episode for each product.
  3. Compute the switch rate, defined as the share of episodes where a different product is applied after the index product.
  4. Run setting-stratified models, or setting-adjusted models if strata are too thin.

The common mistakes are ignoring setting differences and treating high utilization as evidence of effectiveness. Be careful in the other direction too: low application counts can reflect early closure, early discontinuation, or payer constraints, so interpret them alongside outcomes and not in place of them.

Measure applications per episode, switch rate, and results by setting. Platforms that track product utilization across care settings, such as Intellicure Analytics’ Wound Care Industry Dashboard, make these descriptive patterns available as context, though they describe use and do not by themselves establish comparative efficacy.

6. Verify standard-of-care adherence

A skin substitute is meant to be an adjunct to good basic care, and healing attributed to a product may reflect offloading, compression, debridement, or infection control. If adherence differs between product groups, the comparison is confounded by care quality. Payers generally expect documentation of conservative care before advanced products are used, as do the skin substitute LCDs still active in the Novitas, First Coast, and CGS jurisdictions, so this strategy serves both scientific and compliance purposes.

Illustration: exclude episodes lacking documented offloading in DFUs, or documented compression in VLUs, and rerun the comparison. If the product difference narrows or reverses, the original result partly reflected who received adequate baseline care.

  1. Define a conservative-care lookback window before the index date.
  2. Extract offloading and compression fields, and where relevant, debridement and glycemic or vascular assessment.
  3. Flag episodes with infection or ischemia, which affect expected healing and may warrant exclusion or stratification.
  4. Apply the exclusion criteria and report them, including the number of episodes removed at each step.

The mistake is assuming standard of care occurred because it is not documented as absent. Absence of a record is not evidence of adherence, and in routine data the two are indistinguishable. Treat undocumented care as unknown, and run the analysis both ways.

Measure the share of episodes meeting standard-of-care criteria in each product group and the change in results after exclusions. Documentation variability across sites is a real limitation here, and one reason structured wound care documentation matters for research-grade data; a low adherence rate may reflect charting practice and not clinical practice, and the report should say which is more likely given what is known.

7. Report with transparency and reuse for PMCF

A comparison is only as credible as the detail that accompanies it. Reviewers and auditors look for denominators, confidence intervals, how missing data were handled, which analyses were pre-specified, and what the authors concede. Ranking products on small samples, or leaving out the limitations section, is the quickest way to lose that credibility.

Illustration: a pre-registered analysis plan is followed by a peer-reviewed publication, and the same specification is then rerun annually as part of post-market clinical follow-up. Results across years become comparable because the definitions did not drift.

  1. Pre-register the analysis plan, including endpoints, covariates, and exclusions.
  2. Set a minimum n per product and suppress small cells instead of reporting unstable estimates.
  3. Report confidence intervals, not point estimates alone, along with missing-data methods.
  4. Version-control the code, the product dictionary, and the endpoint definitions.
  5. Schedule repeat runs on a defined cadence.

The method also pays off for post-market clinical follow-up (PMCF). Manufacturers selling devices in the European Union must conduct PMCF under the EU Medical Device Regulation, and US PMA products may carry FDA post-approval study conditions. In either case, a locked specification turns each cycle into an update rather than a new project. Intellicure Analytics offers Comparative Effectiveness Studies and PMCF services built around this kind of repeatable design.

Measure whether the minimum n per product was met, whether results reproduce on rerun with identical inputs, and how many reviewer or auditor queries were resolved by the documentation alone. Unresolved queries show where the method statement is thin.

Sequencing the work: what to lock down first

Start with the product-set definition, stratification, and endpoint pre-specification. Every later step depends on them, and errors there cannot be repaired by better statistics downstream. Once those are fixed, add the active-comparator real-world design and transparent reporting, then reuse the same framework for PMCF so that each update is comparable to the last. Setting and standard-of-care checks slot in as sensitivity layers once the core design is stable.

Applied together, these steps will not eliminate confounding in observational data, and a report should never imply that they do. They make the limitations visible and the conclusions proportionate, which is what auditors, payers, and peer reviewers actually test. Learn more about our services.