Mostrando las entradas con la etiqueta GRADE. Mostrar todas las entradas
Mostrando las entradas con la etiqueta GRADE. Mostrar todas las entradas

13 noviembre, 2013

GRADE: Construcción de equivalentes terapéuticos

Fuente: Galo Sanchez
Desde hace unas semanas se ven en prensa y otros medios de comunicación opiniones sobre los equivalentes terapéuticos, pero algunas de ellas están desajustadas porque no están dentro del concepto de “equivalencia terapéutica para…” Añadimos al grupo nominal la preposición “para…” porque no es correcto referirse a “equivalencia terapéutica”, pues lo acertado es decir “equivalencia terapéutica para una o para varias variables”, siendo las más deseables las variables de resultados en salud.

            Para pensar y centrar los debates constructivamente, hemos elaborado una sencilla presentación que muestra gráficamente qué es y cómo se construye y calcula mediante el Margen de No Inferioridad (MNI) clínicamente relevante, y la hemos puesto a disposición de los lectores en evalmed.es, pestaña  VARIOS. También puede hacerse directamente en http://evalmedicamento.weebly.com/8/post/2013/10/tica-para-la-construccin-de-grupos-de-equivalentes-teraputicos.html



24 junio, 2013

Understanding Confidence Intervals

Source: Evidenceinmedicine


My previous post, though not really intended to be focused on p values, led to a long and interesting discussion on the issue by Revere.
Revere commented that it wasn't so clear whether I was writing as a frequentist or a Bayesian. In reality I'm a consumer of biostatistics, not a statistician. For me, the issue is pretty simply to try to know whether a given result I'm seeing is true.
Although I wrote about this in my last post I didn't frame it in exactly this way: the problem with p values is that I want to know how likely it is that the result I'm seeing is correct, while the p value tells me how unlikely it is that I would be seeing a given result if the null hypothesis were true. This is similar to the problem with knowing the sensitivity and specificity of a test, when what you really want to know is the positive predictive value (how likely is it that the patient has the disease now that the test has come back positive?).
The academic medical world started hearing a lot about using confidence intervals in preference to p values 15 or 20 years ago, and I think that led to many doctors concluding that confidence intervals solve this basic p value problem. Confidence intervals have some real benefits compared with p values, but this is not one of them. To look at this, we need to try to understand what a CI means.
Again using the HIV vaccine example from my last post, we can look at the same parameter that had a p value of 0.04, but now examine the point estimate of efficacy (31.2%) and its 95% CI (1.1 to 52.1%). What is it that has a 95% chance of being true, given that CI of 1.1 to 52.1%?
However tricky p value interpretation is for physicians, understanding the meaning of this CI is much worse. I've almost never heard a physician able to correctly interpret the meaning of such a CI when put on the spot with the above question. When trying to teach the interpretation I repeatedly get told there must be a simpler way to communicate the idea. Given my years of failure at this, I have little hope that this post will adequately clarify things for most readers, so if others who teach this have found a way that works to explain it, please write!
All we can really say about that CI is that (excluding any problems with the design or performance of the trial) if we performed the same study 100 times and calculated 95% CIs each time, we would expect that 95 of the 100 CIs formed in this way would include the true point estimate of vaccine efficacy.
Note that just as with p values, this isn't what we want to know: we want to know how likely it is that the true value is inside the particular CI that we are looking at, but that isn't what the CI actually tells us.
However, despite recognizing what the CI really does and does not tell us, consumer of biostatistics that I am, I (and others) approach CIs operationally: we choose to interpret the CI as a range of values with which the data are reasonably compatible, and to interpret values outside the CI as reasonably incompatible with the data. So, other things being equal, I would say that a vaccine efficacy of 5% was compatible with the results of the NEJM study, while a vaccine efficacy of 60% was not. This does not mean that I think the study has excluded the possibility of the vaccine having 60% efficacy, just that this would be unusual under the plays of chance.
This operational definition works as I decide how to write recommendations about whether to administer such a vaccine. If the vaccine truly had an efficacy of 31% I would likely recommend wide use in high risk patients. If the high end of the CI were true (52% efficacy), I might recommend universal vaccination. If the low end were true (1% efficacy) I would probably recommend leaving it on the shelf. Looking at this, I can quickly realize that if this trial were the only information I had about vaccine efficacy then I have inadequately precise data to support whatever recommendation (or set of recommendations) I might want to make about administering HIV vaccine.
If I were using the GRADE scheme for grading such recommendations, I might have started with the assumption that I had high quality evidence from a large randomized trial. But when I realize that I would make different recommendations based on the reasonable values at each end of the CI, I know that I must downgrade the quality of the evidence for such  imprecision. Recommendations for HIV vaccine based on this trial, using the GRADE scheme, would certainly be graded as having no better than moderate quality evidence because of this imprecision.
Instead of some arbitrary definition of whether a trial is large enough or precise enough, using the CI in this way allows me to communicate something important about the quality of the evidence as I grade recommendations. I downgrade for imprecision not because a CI crosses a null effect boundary (like 1.0 for a relative risk) but because the CI crosses a clinical boundary where the appropriate recommendation would change from one side of the boundary to the other.
In this way, the CI is far more useful than a p value. I keep in the back of my mind, though, that the CI doesn't really mean what I'm trying to use it to mean -- it's just that I usually don't have anything better.
I'll write more in the future about how GRADE looks at other types of limitations on the quality of evidence from randomized trials.

10 mayo, 2013

Eficacia de bazedoxifeno en la reducción del riesgo de nuevas fracturas vertebrales en mujeres posmenopáusicas.

File: Osteoporosis-international-179x237.jpg
File: Osteoporosis-international-179x237.jpg (Photo credit: Wikipedia)
Via: Galo Sanchez 
Miguel Ángel Martín de la Nava[1] ha hecho el resumen de una evaluación GRADE del ensayo clínico 300-WW, que hemos puesto a disposición de los interesados en la web de la Oficina (evalmed.es), en la pestaña “FORMACIÓN”, cuyas recomendaciones (válidas para este estudio) se recuadran más abajo.


Estudio 300-WW: Eficacia de bazedoxifeno en la reducción del riesgo de nuevas fracturas vertebrales en mujeres posmenopáusicas.

Silverman SL, Christiansen C, Genant HK. Efficacy of Bazedoxifene in Reducing New Vertebral Fracture Risk in Posmenopausal Women With Osteoporosis: Results From a 3-Year, Randomized, Placebo-, and Active-Controlled Clinical Trial. J Bone Miner Res 2008;23:1923–34.

Los resultados que más importan a las mujeres posmenopáusicas con un diagnóstico de osteoporosis son las fracturas del cuello del fémur (también llamadas de cadera), por la morbimortalidad asociada, y las fracturas vertebrales clínicas.
     El estudio 300-WW pretende evaluar la eficacia y seguridad de bazedoxifeno frente a raloxifeno y frente a placebo en la prevención de fracturas vertebrales y no vertebrales.

RECOMENDACIONES GRADE (VÁLIDAS PARA ESTE ESTUDIO)


Justificación:

A) BENEFICIOS Y RIESGOS AÑADIDOS: Frente a placebo, tanto bazedoxifeno como raloxifeno no muestran beneficios en fracturas de cuello de fémur (llamadas de cadera), fracturas NO vertebrales totales, fracturas vertebrales clínicas, morbimortalidad cardiovascular, ni carcinoma de mama o endometrio. Únicamente muestran beneficio significativo en las fracturas vertebrales totales, si bien a expensas de las “subclínicas”.
                En riesgos añadidos se asocian con más trombosis venosa profunda (aunque sus NND son de magnitud de efecto muy baja), sofocos (mayoritariamente leves o moderados) y calambres en las piernas.

B) INCONVENIENTES: Tomar una pastilla diaria.

C) COSTES: Coste/día: RALO-60: 0,74 €/día (270 €/año) y BAZE-20: 1,23 €/día (449 €/año).
                Acudiendo a los intervalos de confianza del NNT, el coste económico de la prevención de una fractura vertebral (a expensas de las subclínicas) con BAZE-20, en las condiciones de este ensayo clínico, estaría entre 46.048 y 254.363 euros en 3 años, y con RALO-60 entre 27.690 y 152.958 euros en 3 años. No incluimos los costes de la asistencia médica ni de las pruebas diagnósticas, porque excede del objetivo de esta evaluación.


[1] Miguel Ángel Martín de la Nava. Farmacéutico de Atención Primaria. Centro de Salud de Zorita (Cáceres)

29 abril, 2013

Estadisticamente importante pero clinicamente significativo ?

 Source: Galo Sanchez

MISIÓN DE LAS INTERVENCIONES DE EVALUACIÓN, INFORMACIÓN Y FORMACIÓN DEL GRUPO GRADE.

La misión de toda intervención sanitaria es disminuir en una magnitud relevante los riesgos[1] basales graves y moderados de un individuo, sin que como consecuencia de esa intervención se le añada un daño tal que iguale o supere el de su situación inicial. 
            En tanto que nuestras evaluaciones, informaciones y formación son intervenciones sanitarias, nuestra misión específica es que los profesionales sanitarios trasladen los resultados en salud procedentes de la mejor evidencia disponible a su práctica para cumplir con la misión expuesta en el primer párrafo.
            Los medios que utilizamos para esta misión son:
            a) poner en práctica la metodología GRADE para buscar, calcular, reunir y proporcionar los resultados en salud procedentes de la investigación científica en un formato de beneficios y riesgos que justifiquen los inconvenientes y los costes (balance BRIC).
            b) enseñar a los profesionales a calcular por sí mismos en unos pocos segundos la relevancia clínica o poblacional de los resultados que leen en las publicaciones científicas cuando éstas no se los ofrecen, y mejorar con ello su autoeficacia y autoconfianza[2].

[1] No confundir “riesgo” con “factor de riesgo”. Actualmente hay más de cien “factores de riesgo cardiovascular”, que son asociaciones estadísticas entre tales factores y los “riesgos”. Efectivamente, los factores de riesgo son asociaciones estadísticas y no las causas, por lo cual la intervención artificial sobre ellos no significa inequívocamente que disminuirá el riesgo con el que está asociado estadísticamente.
[2] De la razón científica y razón ética que fundamentan la práctica sanitaria, y muy especialmente la práctica clínica, nuestra misión es mejorar la primera.


Documento completo en  PDF

30 junio, 2011

Grading quality of evidence and strength of recommendations for diagnostic tests and strategies

Clinical governance is an aggregation of servi...Image via Wikipedia

  • Rating quality of evidence and strength of recommendations

Grading quality of evidence and strength of recommendations for diagnostic tests and strategies

  1. Holger J Schünemann, professor12, 
  2. Andrew D Oxman, researcher3,
  3. Jan Brozek, research fellow1, 
  4. Paul Glasziou, professor4, 
  5. Roman Jaeschke, clinical professor5, 
  6. Gunn E Vist, researcher3, 
  7. John W Williams Jr, professor6,
  8. Regina Kunz, associate professor7, 
  9. Jonathan Craig, associate professor8,
  10. Victor M Montori, associate professor9, 
  11. Patrick Bossuyt, professor10,
  12. Gordon H Guyatt, professor2
  13.  
  14. for the GRADE Working Group
+Author Affiliations
  1. 1Department of Epidemiology, Italian National Cancer Institute Regina Elena, 00144 Rome, Italy
  2. 2CLARITY Research Group, Department of Clinical Epidemiology and Biostatistics, McMaster University, Hamilton, Ontario, Canada L8N 3Z5
  3. 3Norwegian Knowledge Centre for the Health Services, PO Box 7004, 0130 Oslo, Norway
  4. 4Centre for Evidence-Based Medicine, Department of Primary Health Care, University of Oxford, Oxford OX3 7LF
  5. 5Department of Medicine, McMaster University, 1200 Main Street West, Hamilton, Ontario, Canada L8N 3Z5
  6. 6Department of Medicine, Duke University and Durham VA Medical Center, Durham, NC 27705, USA
  7. 7Basel Institute of Clinical Epidemiology, University Hospital Basel, Hebelstrasse 10, 4031 Basel, Switzerland
  8. 8Screening and Test Evaluation Program, School of Public Health, University of Sydney, Department of Nephrology, Children’s Hospital at Westmead, Sydney, Australia
  9. 9Knowledge and Encounter Research Unit, Department of Medicine, Mayo Clinic College of Medicine, Rochester, MN 55905, USA
  10. 10Department of Clinical Epidemiology, Biostatistics and Bioinformatics, Academic Medical Centre, University of Amsterdam, Amsterdam 1100 DE, Netherlands
  1. Correspondence to: H J Schünemann schuneh@mcmaster.ca
    The GRADE system can be used to grade the quality of evidence and strength of recommendations for diagnostic tests or strategies. This article explains how patient-important outcomes are taken into account in this process

    Summary points

    • As for other interventions, the GRADE approach to grading the quality of evidence and strength of recommendations for diagnostic tests or strategies provides a comprehensive and transparent approach for developing recommendations
    • Cross sectional or cohort studies can provide high quality evidence of test accuracy
    • However, test accuracy is a surrogate for patient-important outcomes, so such studies often provide low quality evidence for recommendations about diagnostic tests, even when the studies do not have serious limitations
    • Inferring from data on accuracy that a diagnostic test or strategy improves patient-important outcomes will require the availability of effective treatment, reduction of test related adverse effects or anxiety, or improvement of patients’ wellbeing from prognostic information
    • Judgments are thus needed to assess the directness of test results in relation to consequences of diagnostic recommendations that are important to patients
    In this fourth article of the five part series, we describe how guideline developers are using GRADE to rate the quality of evidence and move from evidence to a recommendation for diagnostic tests and strategies. Although recommendations on diagnostic testing share the fundamental logic of recommendations on treatment, they present unique challenges. We will describe why guideline panels should be cautious when they use evidence of the accuracy of tests (“test accuracy”) as the basis for recommendations and why evidence of test accuracy often provides low quality evidence for making recommendations.

    Testing makes a variety of contributions to patient care

    Clinicians use tests that are usually referred to as “diagnostic”—including signs and symptoms, imaging, biochemistry, pathology, and psychological testing—for various purposes. 1 These purposes include identifying physiological derangements, establishing prognosis, monitoring illness and response to treatment, and diagnosis. This article …