How Accurate are the Extremely Small P-values Used in Genomic Research: An Evaluation of Numerical Libraries.

Sai Santosh Bangalore,Jelai Wang,David B Allison

Computational Statistics & Data Analysis（2009）

引用 7|浏览0

暂无评分

摘要

In the fields of genomics and high dimensional biology (HDB), massive multiple testing prompts the use of extremely small significance levels. Because tail areas of statistical distributions are needed for hypothesis testing, the accuracy of these areas is important to confidently make scientific judgments. Previous work on accuracy was primarily focused on evaluating professionally written statistical software, like SAS, on the Statistical Reference Datasets (StRD) provided by National Institute of Standards and Technology (NIST) and on the accuracy of tail areas in statistical distributions. The goal of this paper is to provide guidance to investigators, who are developing their own custom scientific software built upon numerical libraries written by others. In specific, we evaluate the accuracy of small tail areas from cumulative distribution functions (CDF) of the Chi-square and t-distribution by comparing several open-source, free, or commercially licensed numerical libraries in Java, C, and R to widely accepted standards of comparison like ELV and DCDFLIB. In our evaluation, the C libraries and R functions are consistently accurate up to six significant digits. Amongst the evaluated Java libraries, Colt is most accurate. These languages and libraries are popular choices among programmers developing scientific software, so the results herein can be useful to programmers in choosing libraries for CDF accuracy.

查看译文

关键词

tail area,small tail area,cdf accuracy,statistical distribution,scientific software,small p-values,java library,statistical software,numerical library,genomic research,scientific judgment,r function,cumulative distribution function,hypothesis test,bioinformatics,multiple testing,biomedical research

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要