<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="other" dtd-version="1.2" xml:lang="en">
    <front>
        <journal-meta>
            <journal-id journal-id-type="pmc">F1000Research</journal-id>
            <journal-title-group>
                <journal-title>F1000Research</journal-title>
            </journal-title-group>
            <issn pub-type="epub">2046-1402</issn>
            <publisher>
                <publisher-name>F1000 Research Limited</publisher-name>
                <publisher-loc>London, UK</publisher-loc>
            </publisher>
        </journal-meta>
        <article-meta>
            <article-id pub-id-type="doi">10.12688/f1000research.2-230.v1</article-id>
            <article-categories>
                <subj-group subj-group-type="heading">
                    <subject>Opinion Article</subject>
                </subj-group>
                <subj-group>
                    <subject>Articles</subject>
                    <subj-group>
                        <subject>Bioinformatics</subject>
                    </subj-group>
                    <subj-group>
                        <subject>Genomics</subject>
                    </subj-group>
                </subj-group>
            </article-categories>
            <title-group>
                <article-title>Progress and challenges in the computational prediction of gene function using networks: 2012-2013 update</article-title>
                <fn-group content-type="pub-status">
                    <fn>
                        <p>[version 1; peer review: 2 approved]</p>
                    </fn>
                </fn-group>
            </title-group>
            <contrib-group>
                <contrib contrib-type="author" corresp="yes">
                    <name>
                        <surname>Pavlidis</surname>
                        <given-names>Paul</given-names>
                    </name>
                    <xref ref-type="corresp" rid="c1">a</xref>
                    <xref ref-type="aff" rid="a1">1</xref>
                </contrib>
                <contrib contrib-type="author" corresp="yes">
                    <name>
                        <surname>Gillis</surname>
                        <given-names>Jesse</given-names>
                    </name>
                    <xref ref-type="corresp" rid="c2">b</xref>
                    <xref ref-type="aff" rid="a2">2</xref>
                </contrib>
                <aff id="a1">
                    <label>1</label>Centre for High-Throughput Biology and Department of Psychiatry, University of British Columbia, Vancouver, V6T1Z4, Canada</aff>
                <aff id="a2">
                    <label>2</label>Stanley Institute for Cognitive Genomics, Cold Spring Harbor Laboratory, Woodbury, NY, 11797, USA</aff>
            </contrib-group>
            <author-notes>
                <corresp id="c1">
                    <label>a</label>
                    <email xlink:href="mailto:paul@chibi.ubc.ca">paul@chibi.ubc.ca</email>
                </corresp>
                <corresp id="c2">
                    <label>b</label>
                    <email xlink:href="mailto:Jgillis@cshl.edu">Jgillis@cshl.edu</email>
                </corresp>
                <fn fn-type="con">
                    <p>PP and JG conceived and wrote the article.</p>
                </fn>
                <fn fn-type="conflict">
                    <p>
                        <bold>Competing interests: </bold>No competing interests were disclosed.</p>
                </fn>
            </author-notes>
            <pub-date pub-type="epub">
                <day>31</day>
                <month>10</month>
                <year>2013</year>
            </pub-date>
            <pub-date pub-type="collection">
                <year>2013</year>
            </pub-date>
            <volume>2</volume>
            <elocation-id>230</elocation-id>
            <history>
                <date date-type="accepted">
                    <day>21</day>
                    <month>10</month>
                    <year>2013</year>
                </date>
            </history>
            <permissions>
                <copyright-statement>Copyright: &#x00a9; 2013 Pavlidis P and Gillis J</copyright-statement>
                <copyright-year>2013</copyright-year>
                <license xlink:href="https://creativecommons.org/licenses/by/3.0/">
                    <license-p>This is an open access article distributed under the terms of the Creative Commons Attribution Licence, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
                </license>
            </permissions>
            <self-uri content-type="pdf" xlink:href="https://f1000research.com/articles/2-230/pdf"/>
            <related-article elocation-id="10.12688/f1000research.1-14.v1" id="related-article-version-109" journal-id="F1000Research" journal-id-type="pmc" related-article-type="companion" vol="1">
                <article-title>Progress and challenges in the computational prediction of gene function using networks</article-title>
                <pub-id pub-id-type="doi">10.12688/f1000research.1-14.v1</pub-id>
            </related-article>
            <abstract>
                <p>In an opinion published in 2012, we reviewed and discussed our studies of how gene network-based guilt-by-association (GBA) is impacted by confounds related to gene multifunctionality. We found such confounds account for a significant part of the GBA signal, and as a result meaningfully evaluating and applying computationally-guided GBA is more challenging than generally appreciated. We proposed that effort currently spent on incrementally improving algorithms would be better spent in identifying the features of data that do yield novel functional insights. We also suggested that part of the problem is the reliance by computational biologists on gold standard annotations such as the Gene Ontology. In the year since, there has been continued heavy activity in GBA-based research, including work that contributes to our understanding of the issues we raised. Here we provide a review of some of the most relevant recent work, or which point to new areas of progress and challenges.</p>
            </abstract>
            <funding-group>
                <funding-statement>PP was supported by NIH Grant GM076990 and salary awards from the Michael Smith Foundation for Health Research and the Canadian Institutes for Health. JG was supported by a grant from T. and V. Stanley.</funding-statement>
            </funding-group>
        </article-meta>
    </front>
    <body>
        <sec>
            <title>Building better networks</title>
            <p>One of the problems with network-based approaches we documented in our previous papers
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;
                    <xref ref-type="bibr" rid="ref-3">3</xref>
                </sup> is their tendency to converge on &#x201c;easy answers&#x201d;, by which we mean picking genes as candidates for a given disease or function simply because they are involved in many diseases (multifunctional) or are prominent in the network (e.g., hubs)
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>. A possible solution would be to tailor the network data to particular contexts. Multifunctional genes would then have less of a dominant role, because fewer of their functions would be relevant to the network, and the network might reflect this. Fortuitously, several studies that improve our understanding of the utility of context-specific networks recently appeared, though they do not address the questions of whether they reduce multifunctionality and node degree biases.</p>
            <p>Guan 
                <italic toggle="yes">et al.</italic> (2012) constructed 107 tissue-specific networks for the laboratory mouse to be used in disease-gene prioritization
                <sup>
                    <xref ref-type="bibr" rid="ref-4">4</xref>
                </sup>. They used a combination of training data from Gene Ontology (GO) and tissue-specific expression signatures to customize their networks before moving to predicting disease candidate genes. The networks are not built from tissue-specific data, but various data used in combination, with each given a weight computed using "tissue-specific gold standards". The cross-validation performance improvement was significant but modest across most tasks (appearing to be approximately 0.03 on top of areas under receiver operating characteristic curves (AUROCs) ranging from 0.7 to 0.8).</p>
            <p>Magger 
                <italic toggle="yes">et al.</italic> (2012) took a different approach, choosing to construct tissue-specific protein interaction networks by down-weighting edges involving genes not expressed in the given tissue
                <sup>
                    <xref ref-type="bibr" rid="ref-5">5</xref>
                </sup>. Their baseline performances are somewhat higher than Guan 
                <italic toggle="yes">et al.</italic> and also show improvement with tissue-specificity (from a mean AUROC of 0.82 to ~0.88). However, the bulk of this performance improvement comes with simply removing genes not expressed in a given tissue. Magger 
                <italic toggle="yes">et al.</italic> provide some evidence that edges involving such genes are the source of prediction errors, as simply down-ranking the genes after analyzing a generic (non-tissue-specific) network was not as effective. While Magger 
                <italic toggle="yes">et al.</italic> primarily restricted themselves to examining disease-gene associations where the causal gene was judged to be tissue-specific, they did examine the full disease-gene data where gains from tissue-specificity were much more modest (approximately AUROC 0.83 to 0.845). In contrast to the specialized task, node removal and attenuation of unexpressed genes performed particularly badly but only at low false positive rates (FPR). At high FPRs, node removal outperformed other methods, suggesting a trade-off between high precision and low precision prediction.</p>
            <p>Piro 
                <italic toggle="yes">et al.</italic> (2012) created co-expression networks from groups of genes expressed in specific mouse brain regions and used these together with other information to predict gene-disease relations from 
                <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/guide/">OMIM</ext-link>
                <sup>
                    <xref ref-type="bibr" rid="ref-6">6</xref>
                </sup>. They report that integrating tissue-specific data substantially raises their candidate disease prioritization performance. However, the final performance does not appear to be better than reported using quite old methods (~AUROC of 0.8 overall). Because they do not present results using a comparable &#x201c;generic&#x201d; network, it is difficult to tell if anything was gained.</p>
            <p>Dowell 
                <italic toggle="yes">et al.</italic> (2013) describe the creation and analysis of a mouse embryonic stem cell (mEPSC) specific gene network, relying on extensive manual curation
                <sup>
                    <xref ref-type="bibr" rid="ref-7">7</xref>
                </sup>. This paper caught our attention in part because Dowell 
                <italic toggle="yes">et al.</italic> acknowledge the potential for node degree bias and other issues, and claim &#x201c;we address many of these potential pitfalls&#x201d;. However, we were unable to identify the evidence that their methods do so; indeed, the focus on network hubs combined with very high performance of negative control data (assembled from datasets excluding mESCs), suggests multifunctional biases may have had a role. They suggest the use of cell-type-specific data should &#x201c;reduce the impact of multi-functional genes&#x201d;, but do not report whether this was indeed the case. This might have been of value in explaining how their network results were more specific, even if performance was not higher than some of the negative controls overall. If context-specific data reduces generic effects, it is of utility even if it yields no improvements in performance as judged by the usual metrics.</p>
            <p>These reports can be considered encouraging, but still leave open the question of whether parsing data into more specific subsets is worthwhile, despite the hopes we expressed last year on this count. The noisiness of biological data may be such that breaking data into smaller bins can cost more in terms of robustness than we gain in terms of specificity. We also note that some earlier approaches combine a wide array of expression data, and treat the data sets as features to be weighted in the prioritization method
                <sup>
                    <xref ref-type="bibr" rid="ref-8">8</xref>,
                    <xref ref-type="bibr" rid="ref-9">9</xref>
                </sup>. Thus information such as &#x201c;gene A is expressed in tissue X&#x201d; may have been implicitly used. Choosing &#x201c;tissue-specific functions&#x201d; to assess such approaches is another challenge, and it is unknown if multifunctionality effects are reduced. None of this eliminates the possibility that network specificity provides crucial value, but more data are required.</p>
        </sec>
        <sec>
            <title>Using better controls</title>
            <p>The GBA studies discussed in the previous section did not take the opportunity to test the effect of multifunctionality or node degree bias, despite this control being easy to perform. The clearest attempt to control for multifunctionality and node degree of which we are aware is reported by Singh-Blom and colleagues in characterizing their prediction tool, CATAPAULT. Singh-Blom 
                <italic toggle="yes">et al.</italic> conducted an analysis of disease and drug-target genes with a variety of networks and algorithms
                <sup>
                    <xref ref-type="bibr" rid="ref-10">10</xref>
                </sup>. They report that a ranking of genes by multifunctionality (which they refer to as &#x201c;degree&#x201d;) performs poorly but not negligibly as a predictor, outperforming some of the methods they tested in cross-validation.</p>
            <p>Before we comment further, there are some nuances to how Singh-Blom 
                <italic toggle="yes">et al.</italic> (2013) use the multifunctionality ranking, compared to how we did. First, to avoid confusion the multifunctionality ranking (or a node degree ranking) should not be treated as a &#x201c;method&#x201d; for prediction, as it is referred to by Singh-Blom 
                <italic toggle="yes">et al.</italic> It should be considered a null. In addition, the multifunctionality ranking is expected to &#x201c;perform best&#x201d; when performance is measured using ROC curves; Singh-Blom 
                <italic toggle="yes">et al.</italic> use something more akin to precision-recall, which tends to obscure the influence of multifunctionality (and node degree), while being more heavily influenced by critical edges. Finally, multifunctionality ranking may be a too-stringent control (when using ROC) because it is literally optimized and performs better than many real algorithms on actual data
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>; rather, the correlation of the &#x201c;real&#x201d; prediction results with multifunctionality ranking is often a more helpful measure.</p>
            <p>In any case, the fact that the multifunctionality ranking yields even modest performance in the evaluation scheme of Singh-Blom 
                <italic toggle="yes">et al.</italic> hints that this ranking is also correlated with network node degree (as explained in our work
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>), and furthermore that such effects will have a strong impact on their results. Accordingly, Singh-Blom 
                <italic toggle="yes">et al.</italic> report that highly prioritized genes tend to have high network node degree. They also report that of the top 10 candidates for eight diseases, almost all were shared among two or more of the diseases, confirming our finding that GBA too often yields &#x201c;generic&#x201d; predictions, not function-specific ones. These results, using different algorithms, networks and evaluation metrics than us, provide strong independent support of our claims.</p>
            <p>While Singh-Blom 
                <italic toggle="yes">et al.</italic> confirmed some of our key findings, we may differ with them in interpretation, as they argue that the results are unproblematic. We agree with them that non-specific predictions could be correct, but this does not absolve concern about how their methods are actually operating. For example, INSR (insulin receptor) was predicted by CATAPAULT as a leukemia candidate gene, in addition for other diseases (including other cancers and diabetes). We are led to suspect this is at least partly explained by the high node degree of INSR, and not specific &#x201c;guilt by association&#x201d; of INSR with known leukemia genes. Can one find literature that connects insulin receptors and leukemia? Of course, since both are highly studied (one can find papers linking insulin receptors or cancer to many things) and the metabolism of cancer cells is of interest from a therapeutic standpoint. We also note that TP53 was predicted to be a diabetes-related gene by CATAPAULT. It remains possible that INSR is a 
                <italic toggle="yes">bona fide</italic> leukemia gene; regardless, we strongly believe that a biologist wanting to use the output of CATAPAULT would also want to know about the specificity of the predictions.</p>
            <p>We hope that other researchers interested in why their methods work take the step of attempting to control for &#x201c;generic&#x201d; results. Otherwise methodological performance is open to profound misinterpretation as to true utility. This is true even in the case where authors are clearly aware of the potential for problems. For example, Zuberi 
                <italic toggle="yes">et al.</italic> (2013) report that the GeneMANIA edge weight normalization &#x201c;helps to reduce the impact that the pleiotropy of high degree nodes has on functional predictions&#x201d;
                <sup>
                    <xref ref-type="bibr" rid="ref-9">9</xref>
                </sup>. We showed previously GeneMANIA&#x2019;s results (with the normalization) are strongly affected by node degree, and indeed a substantial fraction of performance as measured by ROC curves could be explained by node degree effects
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>. Likewise, Verbeke 
                <italic toggle="yes">et al.</italic> (2013) describe a gene prioritization method based on local networks that they &#x201c;assume&#x201d; reduces the effect of hubs, but provide no direct test
                <sup>
                    <xref ref-type="bibr" rid="ref-11">11</xref>
                </sup>. As we have documented
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>, various attempts to modify networks to reduce extremes of node degree at best hide the problems from detection. We are open to the possibility that the approach of Verbeke 
                <italic toggle="yes">et al.</italic> has the desired effect.</p>
        </sec>
        <sec>
            <title>Finding better algorithms</title>
            <p>In the last year, there have been several interesting evaluations of gene function prediction methods. Our interest lies less in which method does best than in what these evaluations expose about the state of the field as a whole.</p>
            <p>B&#x00f6;rnigen 
                <italic toggle="yes">et al.</italic> (2012) performed a comparison of eight disease prioritization tools on a set of 42 disease genes
                <sup>
                    <xref ref-type="bibr" rid="ref-12">12</xref>
                </sup>. The task was to prioritize the correct candidate, given whatever input the method requires (typically involving definition of a training set of genes already associated with the target function, and often a list of ~100 candidates in a genomic interval as starting points rather than a genome-wide list). The results were evaluated with ROC curves and with true positive rates at a given threshold. The authors&#x2019; method, Endeavour
                <sup>
                    <xref ref-type="bibr" rid="ref-13">13</xref>
                </sup>, was among the top evaluation performers. No evaluation of the impact of multifunctionality was undertaken, but our experience suggests that multifunctional genes tend to be prioritized by these types of methods
                <sup>
                    <xref ref-type="bibr" rid="ref-14">14</xref>
                </sup>. The problem of multifunctionality biasing prioritizations may be at least partly due to the difficulty of obtaining less biased training data. The &#x201c;known genes&#x201d; are often going to be biased towards highly-studied genes which do not form a sufficiently specific starting point for making functionally specific predictions. Regardless, B&#x00f6;rnigen 
                <italic toggle="yes">et al.</italic> were unable to clearly distinguish a best or poorest method, and the reasons for differences were not identified; it was speculated that differences in the underlying data used were important.</p>
            <p>The more ambitious Critical Assessment of Functional Annotation (CAFA)
                <sup>
                    <xref ref-type="bibr" rid="ref-15">15</xref>
                </sup> was set up in a model very similar to the (now discontinued) function prediction component of CASP6 and CASP7
                <sup>
                    <xref ref-type="bibr" rid="ref-16">16</xref>,
                    <xref ref-type="bibr" rid="ref-17">17</xref>
                </sup>. Participating groups predicted GO annotations for poorly-annotated proteins, followed by a waiting period during which some of the targets happened to be annotated by GO curators. The submitted algorithms were then assessed for correctness, relying primarily on a novel gene-centric metric that allowed partial credit for predicting a &#x201c;similar&#x201d; term, based on proximity in the GO term graph. Unfortunately, it emerged that this metric led to many methods (including BLAST) being outperformed by a na&#x00ef;ve ranking of functions by prevalence (e.g. simply predict functions which are common overall; this approach ranks third or fourth in molecular function prediction), leading the organizers to exclude some results
                <sup>
                    <xref ref-type="bibr" rid="ref-15">15</xref>
                </sup>. Radivojac 
                <italic toggle="yes">et al.</italic> (2013) concluded that simple sequence analysis methods such as BLAST perform poorly, while more sophisticated methods based on integrating diverse data types are a substantial boon. However, in our separate assessment of a substantial portion of the CAFA data, we found that by more conventional metrics BLAST was among the top performers
                <sup>
                    <xref ref-type="bibr" rid="ref-18">18</xref>
                </sup>.</p>
            <p>Our concerns about how function prediction works are further supported by a closer inspection of the methods that did well in CAFA. The best performing of them frequently have embedded in them aspects of the naive scoring method; that is, they successfully use knowledge of GO structure and term prevalence. In combination with the gene-centric evaluation metric, this creates a misleading impression of predictive power, in much the same way that a cold-reading mentalist can exploit the prevalence of names and medical conditions to impress a gullible audience. It was also possible to benefit from existing annotations for the targets (hot reading
                <sup>
                    <xref ref-type="bibr" rid="ref-18">18</xref>
                </sup>). While successful in the narrow confines of the assessment, if put to use such approaches would only serve to increase the already strong biases in GO. In other words, in a sense CAFA turned out to be less about predicting gene function than about predicting which proteins would be annotated by GO curators and with which terms.</p>
            <p>One of the major issues faced by CAFA is how to operationalize &#x201c;function&#x201d;. They (understandably) pass the buck on this issue and take the GO as an appropriate way to define function, with a number of consequent problems. In contrast, DREAM
                <sup>
                    <xref ref-type="bibr" rid="ref-19">19</xref>
                </sup> is a set of critical assessments motivated, like CAFA, as a means of understanding and improving inference methods (particularly network based ones), but DREAM largely focuses on more specific problems with associated datasets. One DREAM assessment we found interesting (even though it is not strictly about function prediction) focused on breast cancer survival analysis, with the goal of using molecular data (expression and genomic copy number) to improve prediction beyond that provided by clinical features. While a number of performance comparisons were made, one result was that the baseline method - simple Cox regression on clinical features &#x2013; outperforms most methods across most conditions, even when they include the use of the molecular data. In only 10 out of 28 submissions were models incorporating molecular feature data with clinical able to outperform the baseline clinical predictor. On the one hand, this is encouraging: molecular data may be able to contribute something. On the other hand, the very best method using clinical data only was very close in performance to the best performing method using combined data. As is typically the case in machine learning, ensemble methods performed well (this would also have been true in CAFA), although investigator-based choices also appear to have been critical, since the class of purely automated methods performed particularly badly. In addition, the control of incorporating random gene signatures (or generic/multifunctional ones) with clinical data appears not to have been attempted (as might be suggested by previous research
                <sup>
                    <xref ref-type="bibr" rid="ref-20">20</xref>
                </sup>), with permuted case labelling serving as the negative control instead. This leaves open the possibility that to the extent molecular data is of any predictive value at all, it does not provide us with any guidance as to molecular mechanism (by singling out relevant subsets of genes).</p>
            <p>The last assessment we consider is the Critical Assessment of Genome Interpretation (CAGI), which focuses on using sequence data to predict elements of clinical or molecular phenotypes. While the results have not been published formally, the data presented on the CAGI web site are informative (
                <ext-link ext-link-type="uri" xlink:href="https://genomeinterpretation.org/">https://genomeinterpretation.org/</ext-link>). The issue of appropriate controls again rears its head. For example, in the 2011 &#x201c;personal genome project&#x201d; assessment of phenotype prediction, the top performing submission appears to have primarily obtained performance by predicting that rare phenotypes would not occur (&#x201c;due to predicting absence of rare characteristics&#x201d;). In another competition, ROC curves for predicting Crohn&#x2019;s disease from exome data appear to show close to half of teams performing below random (although not significantly so, apparently due to low sample size).</p>
            <p>Why are critical assessments done? An admirably thoughtful discussion of algorithm comparisons noted that most scientists read new papers thinking &#x201c;well, of course they say their method is better, but&#x2026;&#x201d;
                <sup>
                    <xref ref-type="bibr" rid="ref-21">21</xref>
                </sup>. In part, critical assessments were intended to solve this problem: to help us move forward by making truly representative comparisons. It is not clear this is what is happening for function prediction assessments. Instead, we now have a system where researchers agree to participate and organizers have an obligation not to embarrass them. Thus, often only the top-performing methods are discussed, and the organizers have the same capacity to tweak as the original algorithm developers would have, and many of the same incentives. Discovering that most methods perform quite badly should be headline news, but could reduce enthusiasm for participation to the point of killing off future assessments. This is the usual problem of negative results, but scaled up to apply to the whole field through &#x201c;consensus&#x201d;. Perhaps publications of critical assessment should devote equal space to characterizing why methods failed; DREAM&#x2019;s characterization of the poor performance of molecular data offers a toehold on this issue.</p>
            <p>To summarize this section, because there are decreasing returns in tweaking methods, in our view comparisons of algorithms are less important than asking if they work at all and if so, how. Unfortunately this is often very difficult to discern from most of the work that has emerged in the last year, which often vary data as well as algorithms, and do not provide enough information to judge potential drivers of performance such as multifunctionality effects.</p>
        </sec>
        <sec>
            <title>Using prior knowledge</title>
            <p>As we noted last year, gene function prediction shouldn&#x2019;t simply reduce to information retrieval, at least not unwittingly. Organizing existing knowledge and finding overlaps is useful, but is not the principal motivation of network-based methods, which are intended to find novel features in rich data. One way of drawing a distinction is that information retrieval GBA does not as readily suggest novel experiments. Normally, using some experimental feature to draw a functional conclusion suggests that one should try perturbing that experimental feature and observing the result; this will seem redundant if the feature is purely a property of the way the data was explicitly organized. However, the influence of prior knowledge is often hard to discern in the output of prediction methods, so information retrieval can masquerade as 
                <italic toggle="yes">de novo</italic> function prediction.</p>
            <p>It is important to realize that methods motivated by information retrieval are still forms of GBA, and are subject to the same potential problems. For example, Hoehndorf 
                <italic toggle="yes">et al.</italic> (2013) created a network of genes based on semantic similarity of phenotypes of genetic diseases and animal models of diseases (
                <ext-link ext-link-type="uri" xlink:href="http://phenomebrowser.net/">PhenomeNet</ext-link>)
                <sup>
                    <xref ref-type="bibr" rid="ref-22">22</xref>
                </sup>. They then use sets of genetic disease genes from various human databases, and their orthologs in mouse, to evaluate the relevance of this network for identifying gene-disease associations. For example, they rank mouse genes by the similarity of their mutant phenotypes to a target human disease&#x2019;s phenotypes. They claim their approach is not GBA because &#x201c;it does not require prior knowledge of the genetic basis of diseases for its predictions&#x201d;. This is incorrect: because their method uses associations (semantic similarity) and infers &#x201c;guilt&#x201d; (involvement of a gene in a disease) based on this, it is obviously GBA, albeit a simple one where the prediction algorithm is a simple ranking of nodes by similarity. The authors may have been hoping that they don&#x2019;t need to worry about node degree effects and multifunctionality, but we disagree. Using semantic similarity to identify diseases that resemble mouse models seems reasonable; using this to predict disease genes is most definitely GBA and suffers all the same potential pitfalls (and then some).</p>
            <p>To see why, we note that the nodes in the network used by Hoehndorf 
                <italic toggle="yes">et al.</italic> can be regarded as the set of both genes and diseases/phenotypes with edges indicating high semantic similarity across phenotypes. A disease node then provides the training data (a set of associated genes), and nearby gene nodes are the predicted relevant genes. By taking the gene-centred data (mutants, etc.) and treating it as equivalent to disease the authors are incorporating a hypothesis in addition to GBA, not instead of it. That is, a disease is treated conceptually as if it was gene-like. Consider that the same model should work if we were trying to predict effects of mutations from other known mutation effects through cross-validation, which would then be GBA (but also including an information retrieval component).</p>
            <p>The use of various flavors of annotation similarity to build or influence networks is already endemic in function prediction, as we noted previously. A recent example is the work of Youngs 
                <italic toggle="yes">et al.</italic> (2013), who use information on GO annotations to compute priors for function prediction
                <sup>
                    <xref ref-type="bibr" rid="ref-23">23</xref>
                </sup>. In this manner, the likelihood that a gene is predicted to be annotated with a certain GO term is influenced by whether its other annotated GO terms tend to co-occur with desired GO term. This method, which is influenced by early work
                <sup>
                    <xref ref-type="bibr" rid="ref-24">24</xref>
                </sup> performs strongly in cross-validation, but we see two issues. The first is that, once again, the authors claim their evaluation approach addresses biases we have reported, without providing evidence. Second, it treats annotation biases as something to exploit (somewhat like Singh-Blom 
                <italic toggle="yes">et al.</italic>), which we regard as a shaky proposition when it comes to predicting gene function, as opposed to performing information retrieval.</p>
            <p>One interesting feature of the attempts to improve networks, methods, and priors in GBA is that researchers in each area can take the other area to be a gold standard. Thus, researchers focusing on protein interaction networks may use GO to obtain &#x201c;better&#x201d; interaction data
                <sup>
                    <xref ref-type="bibr" rid="ref-25">25</xref>
                </sup>. Conversely, researchers wish to treat the network as a gold standard to improve GO
                <sup>
                    <xref ref-type="bibr" rid="ref-26">26</xref>,
                    <xref ref-type="bibr" rid="ref-27">27</xref>
                </sup>. In the meantime, algorithm developers treat both networks and annotations as a gold standard when comparing methods. In all these cases, researchers are performing what we would call &#x201c;GBA&#x201d; and working with the alignment between how genes form groups (using some method) as characterized in data and how those genes are grouped by prior annotations. The fact that there is some form of alignment is repeatedly rediscovered. A problem we perceive is that the duality of the gold-standards is increasingly blurring the lines between predictions and data. For example, Dutkowski 
                <italic toggle="yes">et al.</italic> (2013), when benchmarking their method against GO, initially used as input some networks that were influenced by data from GO (e.g., YeastNet
                <sup>
                    <xref ref-type="bibr" rid="ref-25">25</xref>
                </sup>), and so had to perform separate experiments to remove this confound. Similarly, the work of Magger 
                <italic toggle="yes">et al.</italic> discussed above used data on disease gene expression patterns
                <sup>
                    <xref ref-type="bibr" rid="ref-28">28</xref>
                </sup> that were derived in part from the same protein interaction data that Magger 
                <italic toggle="yes">et al.</italic> then use to perform tissue-specific predictions, though the implications of this are unclear. Recently we documented how protein interactions and gene ontology annotations are in many cases derived from the same publications
                <sup>
                    <xref ref-type="bibr" rid="ref-29">29</xref>
                </sup>. Data resources used in genomics are becoming more intertwined, so ever greater care is required to avoid contaminating computational experiments with unwanted biases.</p>
        </sec>
        <sec>
            <title>GBA success stories?</title>
            <p>Guilt by association is widely agreed to be a valid method for investigating gene function. As mentioned, our concerns largely have to do with how GBA is performed and evaluated computationally (though the biases in existing knowledge could have impact on GBA even when it is conducted by hand). We also want to know, when GBA does work, is it because of &#x201c;generic features&#x201d; such as node degree, or are the GBA methods working the way most computational biologists hope they are working, which is inferring specific things about a gene based on specific features of its network neighbors. It is therefore of great interest to us and the rest of the computational GBA field to see use of GBA &#x201c;in the wild&#x201d;.</p>
            <p>Our review of the literature reveals different stories for disease gene prioritization and for other function prediction tasks. It is uncommon to see papers that report using computational GBA as an important means of identifying genes with a desired function (ignoring the role of sequence similarity, which is no doubt the most-used GBA method; methods that use more complex network-based approaches are our focus here). In contrast, genetics researchers faced with a genomic interval or a set of candidates seem to more readily turn to prioritization tools for assistance. This may be because the task of prioritizing a few genes (often 10&#x2013;20) is simpler than prioritizing the entire genome, or that the task of identifying a disease gene is more clearly defined.</p>
            <p>To take a well-known method as example of how algorithms are used in practice, 
                <ext-link ext-link-type="uri" xlink:href="http://www.genemania.org/">GeneMANIA</ext-link> is implemented in a web-based tool described by its developers as a gene recommender system &#x2013; essentially, using GBA
                <sup>
                    <xref ref-type="bibr" rid="ref-9">9</xref>
                </sup>. A survey of recent citations suggests that the function prediction aspect is not the focus of most users of GeneMANIA
                <sup>
                    <xref ref-type="bibr" rid="ref-30">30</xref>&#x2013;
                    <xref ref-type="bibr" rid="ref-34">34</xref>
                </sup>; in some cases users express an interest in predicting interactions, but not functions
                <sup>
                    <xref ref-type="bibr" rid="ref-35">35</xref>
                </sup>. This may be because GeneMANIA&#x2019;s tools do not operate on functions as usually defined in GBA settings; they take as an input a set of genes chosen by the user, and show a gene network with nodes selected using the GeneMANIA algorithm. In this way it is very similar to 
                <ext-link ext-link-type="uri" xlink:href="http://string-db.org/">STRING</ext-link>
                <sup>
                    <xref ref-type="bibr" rid="ref-36">36</xref>
                </sup> and 
                <ext-link ext-link-type="uri" xlink:href="http://www.functionalnet.org/humannet/about.html">HumanNet</ext-link>
                <sup>
                    <xref ref-type="bibr" rid="ref-37">37</xref>
                </sup>, which generate gene networks, which are adjusted by their agreement with inputs such as GO. We suspect some users of GeneMANIA, STRING and HumanNet do not realize that they are looking at the output of a guilt-by-association-influenced approach.</p>
            <p>This is not to say that computational methods are not directly used successfully for function prediction. But even in such cases, it is often very difficult to determine how exactly GBA worked. Users of GBA are justifiably relatively uninterested in how they get to an answer, only that it is correct and leads to a new set of testable hypotheses or insights. Thus GBA success stories tend to be somewhat light on details and heavy on ad hoc aspects. We briefly mentioned several success stories in our original commentary. Some additional examples have since come to light and bear discussion.</p>
            <p>Tacutu 
                <italic toggle="yes">et al.</italic> (2012) provide a valuable rare large-scale assessment of a computational GBA task
                <sup>
                    <xref ref-type="bibr" rid="ref-38">38</xref>
                </sup>. They wished to predict genes involved in regulating the longevity of 
                <italic toggle="yes">Caenorhabditis elegans</italic>. Their input is a set of 205 known longevity-associated genes (LAGs) in a worm protein interaction network of 871 genes. A similar network was constructed for human orthologs. They then made predictions in a simple way, considering any gene that was a network neighbor of known LAGs (or orthologs), limited to the subset of candidates that are required for development (essential genes). This yielded 500 candidates of which 374 were tested. They report 19 of these validated, a success rate of 5%, compared to a rate of 2.4% based only on genes critical for development. It is notable that most of the predictive information came from exploiting the prior knowledge that essential genes were good candidates: the success rate went up five-fold from ~0.5% (for genome-wide screens) to 2.4% for essential genes, whereas adding the protein interaction data increased success by another two-fold. Tacutu 
                <italic toggle="yes">et al.</italic> do not report any information on potential node degree effects, but obviously the largest number of candidates would have to come from the highest node-degree LAGs given the simplicity of the method. We suspect that many GBA experts would want to know how it would have worked with &#x201c;better data&#x201d; and a &#x201c;more sophisticated algorithm&#x201d;.</p>
            <p>Putnam 
                <italic toggle="yes">et al.</italic> (2012) sought to identify yeast genes suppressing genomic instability
                <sup>
                    <xref ref-type="bibr" rid="ref-39">39</xref>
                </sup>. Their initial input was 75 genes already known to be involved in genomic instability, and 928 genes in which mutations cause sensitivity to DNA damage. This set was expanded by choosing genes which have similar profiles of genetic interactions &#x2013; this appears to be their key GBA step. The final set of genes numbered 1041, which were further prioritized in a second GBA-like step, based on their genetic interaction profiles with known DNA damage response genes, and by apparently manual selection, leading to selection of 87 genes for experimental follow-up. Of these, 40% had a detectable effect on genomic stability when mutated. This is impressive, but it is difficult to quantify the contribution of computation versus manual selection, and relatively few candidates were entirely novel. The authors speculate that their success rates might have been even higher if they had genetic interaction from a more relevant phenotype. For example, many genes that clustered with known DNA damage-related genes &#x2013; and which thus looked &#x201c;guilty&#x201d; - failed to validate. As in the case with Tacutu 
                <italic toggle="yes">et al.</italic>, we note the relatively simple data used and the simple approach, combined with hand-tuning.</p>
            <p>We have also reviewed some recent applications of disease gene prioritization tools, which as we comment above are seemingly used more commonly than function prediction tools (even though they are conceptually similar). We are struck by two trends. First, many (perhaps most) papers that apply such methods make no strong conclusion as to whether they have found the right gene
                <sup>
                    <xref ref-type="bibr" rid="ref-40">40</xref>&#x2013;
                    <xref ref-type="bibr" rid="ref-48">48</xref>
                </sup>. That is, the results are treated a bit like GO enrichment analyses: as suggestive or exploratory. Second, some papers that report prioritization tool results supplant them with more precise or manually-identified information, such as the existence of an orthologous mouse mutant that has similar phenotypic features
                <sup>
                    <xref ref-type="bibr" rid="ref-49">49</xref>&#x2013;
                    <xref ref-type="bibr" rid="ref-52">52</xref>
                </sup>.</p>
            <p>We note a few trends from these reports. Black-box application of existing prioritization methods played at best a supporting role. The use of custom methods for creating initial target sets were important, sometimes based on experiments under the investigator&#x2019;s control, rather than existing annotations from public databases. Data that was specific to the biology was deemed important: using generic data is a fallback. Not surprisingly, even with these ingredients, success in converting computationally-prioritized genes into documented hits is far from guaranteed. And while these examples bolster the claim that GBA can work, exactly how they are working with regards to multifunctionality bias is still left unclear.</p>
        </sec>
        <sec sec-type="conclusions">
            <title>Conclusions</title>
            <p>A theme that emerges from our review is brought out by the difference between the practices of computational biologists and those who actually use function prediction tools (loosely defined). These differences should come as no real surprise, but it has important implications that we feel are not being attended to sufficiently. It may be that biologists are happy with high-quality information retrieval tools, and are not actually very interested in function prediction at all. That creates a difficulty for those who are interested in predicting function, who feel compelled to develop new methods to do so, and who want their tools to be used by others. Such practitioners are left to test their methods on the GO, which we are increasingly certain is a waste of time, in the sense that it isn&#x2019;t realistic, it isn&#x2019;t what interests biologists, and it is easily confounded with the data used for prediction. It remains difficult to tell when methods are actually doing something useful, because evaluations have been weak, and the &#x201c;in the wild&#x201d; uses are obfuscated by various hand-tunings, publication bias, or inadvertent cherry-picking.</p>
            <p>We regard these as major issues, but this is a far cry from disagreeing with functional inference overall. Some authors appear to have interpreted our papers as concluding that GBA is useless
                <sup>
                    <xref ref-type="bibr" rid="ref-53">53</xref>,
                    <xref ref-type="bibr" rid="ref-54">54</xref>
                </sup>, but this is too broad a brush. We fully believe in the GBA principle, and computational methods can be useful. The difficulty is in telling 
                <italic toggle="yes">when</italic> and 
                <italic toggle="yes">how</italic> they are working. The concern about &#x201c;when&#x201d; is summarized by our findings that cross-validation analysis can be very misleading &#x2013; to the point of being potentially irrelevant - for predicting future performance
                <sup>
                    <xref ref-type="bibr" rid="ref-2">2</xref>
                </sup>. The concern about &#x201c;how&#x201d; is reflected in our demonstrations that gene multifunctionality and node degree effects are often more important in determining the outcome of a GBA analysis than details about the connections in the network
                <sup>
                    <xref ref-type="bibr" rid="ref-1">1</xref>
                </sup>. These two realizations should affect practice, but they do not mean that the predictions one makes are always incorrect.</p>
        </sec>
    </body>
    <back>
        <ref-list>
            <ref id="ref-1">
                <label>1</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Gillis</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Pavlidis</surname>
                            <given-names>P</given-names>
                        </name>
					</person-group>:
                    <article-title>The impact of multifunctional genes on "guilt by association" analysis.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2011</year>;<volume>6</volume>(<issue>2</issue>):<fpage>e17258</fpage>.
                    <pub-id pub-id-type="pmid">21364756</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0017258</pub-id>
                    <pub-id pub-id-type="pmcid">3041792</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-2">
                <label>2</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Gillis</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Pavlidis</surname>
                            <given-names>P</given-names>
                        </name>
					</person-group>:
                    <article-title>'Guilt by association&#x2019; is the exception rather than the rule in gene networks.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS Comput Biol.</italic>
					</source>
                    <year>2012</year>;<volume>8</volume>(<issue>3</issue>):<fpage>e1002444</fpage>.
                    <pub-id pub-id-type="pmid">22479173</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002444</pub-id>
                    <pub-id pub-id-type="pmcid">3315453</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-3">
                <label>3</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Pavlidis</surname>
                            <given-names>P</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gillis</surname>
                            <given-names>J</given-names>
                        </name>
					</person-group>:
                    <article-title>Progress and challenges in the computational prediction of gene function using networks.</article-title>
                    <source>
						
                        <italic toggle="yes">F1000 Res.</italic>
					</source>
                    <year>2012</year>;<volume>1</volume>:<fpage>1</fpage>&#x2013;<lpage>14</lpage>.
                    <pub-id pub-id-type="pmid">23936626</pub-id>
                    <pub-id pub-id-type="doi">10.12688/f1000research.1-14.v1</pub-id>
                    <pub-id pub-id-type="pmcid">3653610</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-4">
                <label>4</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Guan</surname>
                            <given-names>Y</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gorenshteyn</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Burmeister</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Tissue-specific functional networks for prioritizing phenotype and disease genes.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS Comput Biol.</italic>
					</source>
                    <year>2012</year>;<volume>8</volume>(<issue>9</issue>):<fpage>e1002694</fpage>.
                    <pub-id pub-id-type="pmid">23028291</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002694</pub-id>
                    <pub-id pub-id-type="pmcid">3459891</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-5">
                <label>5</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Magger</surname>
                            <given-names>O</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Waldman</surname>
                            <given-names>YY</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Ruppin</surname>
                            <given-names>E</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Enhancing the prioritization of disease-causing genes through tissue specific protein interaction networks.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS Comput Biol.</italic>
					</source>
                    <year>2012</year>;<volume>8</volume>(<issue>9</issue>):<fpage>e1002690</fpage>.
                    <pub-id pub-id-type="pmid">23028288</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002690</pub-id>
                    <pub-id pub-id-type="pmcid">3459874</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-6">
                <label>6</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Piro</surname>
                            <given-names>RM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Molineris</surname>
                            <given-names>I</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Di Cunto</surname>
                            <given-names>F</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Disease-gene discovery by integration of 3D gene expression and transcription factor binding affinities.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>29</volume>(<issue>4</issue>):<fpage>468</fpage>&#x2013;<lpage>475</lpage>.
                    <pub-id pub-id-type="pmid">23267172</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/bts720</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-7">
                <label>7</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Dowell</surname>
                            <given-names>KG</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Simons</surname>
                            <given-names>AK</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Wang</surname>
                            <given-names>ZZ</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Cell-type-specific predictive network yields novel insights into mouse embryonic stem cell self-renewal and cell fate.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>(<issue>2</issue>):<fpage>e56810</fpage>.
                    <pub-id pub-id-type="pmid">23468881</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0056810</pub-id>
                    <pub-id pub-id-type="pmcid">3585227</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-8">
                <label>8</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Hibbs</surname>
                            <given-names>MA</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Hess</surname>
                            <given-names>DC</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Myers</surname>
                            <given-names>CL</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Exploring the functional landscape of gene expression: directed search of large microarray compendia.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2007</year>;<volume>23</volume>(<issue>20</issue>):<fpage>2692</fpage>&#x2013;<lpage>2699</lpage>.
                    <pub-id pub-id-type="pmid">17724061</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/btm403</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-9">
                <label>9</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Zuberi</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Franz</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Rodriguez</surname>
                            <given-names>H</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>GeneMANIA prediction server 2013 update.</article-title>
                    <source>
						
                        <italic toggle="yes">Nucleic Acids Res.</italic>
					</source>
                    <year>2013</year>;<volume>41</volume>(<issue>Web Server issue</issue>):<fpage>W115</fpage>&#x2013;<lpage>W122</lpage>.
                    <pub-id pub-id-type="pmid">23794635</pub-id>
                    <pub-id pub-id-type="doi">10.1093/nar/gkt533</pub-id>
                    <pub-id pub-id-type="pmcid">3692113</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-10">
                <label>10</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Singh-Blom</surname>
                            <given-names>UM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Natarajan</surname>
                            <given-names>N</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tewari</surname>
                            <given-names>A</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Prediction and validation of gene-disease associations using methods inspired by social network analyses.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>(<issue>5</issue>):<fpage>e58977</fpage>.
                    <pub-id pub-id-type="pmid">23650495</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0058977</pub-id>
                    <pub-id pub-id-type="pmcid">3641094</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-11">
                <label>11</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Verbeke</surname>
                            <given-names>LP</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Cloots</surname>
                            <given-names>L</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Demeester</surname>
                            <given-names>P</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>EPSILON: an eQTL prioritization framework using similarity measures derived from local networks.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>29</volume>(<issue>10</issue>):<fpage>1308</fpage>&#x2013;<lpage>1316</lpage>.
                    <pub-id pub-id-type="pmid">23595663</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/btt142</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-12">
                <label>12</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>B&#x00f6;rnigen</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tranchevent</surname>
                            <given-names>LC</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Bonachela-Capdevila</surname>
                            <given-names>F</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>An unbiased evaluation of gene prioritization tools.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2012</year>;<volume>28</volume>(<issue>23</issue>):<fpage>3081</fpage>&#x2013;<lpage>8</lpage>.
                    <pub-id pub-id-type="pmid">23047555</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/bts581</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-13">
                <label>13</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Tranchevent</surname>
                            <given-names>LC</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Barriot</surname>
                            <given-names>R</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Yu</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>ENDEAVOUR update: a web resource for gene prioritization in multiple species.</article-title>
                    <source>
						
                        <italic toggle="yes">Nucleic Acids Res.</italic>
					</source>
                    <year>2008</year>;<volume>36</volume>(<issue>Web Server issue</issue>):<fpage>W377</fpage>&#x2013;<lpage>W384</lpage>.
                    <pub-id pub-id-type="pmid">18508807</pub-id>
                    <pub-id pub-id-type="doi">10.1093/nar/gkn325</pub-id>
                    <pub-id pub-id-type="pmcid">2447805</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-14">
                <label>14</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Qiao</surname>
                            <given-names>Y</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Harvard</surname>
                            <given-names>C</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tyson</surname>
                            <given-names>C</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Outcome of array CGH analysis for 255 subjects with intellectual disability and search for candidate genes using bioinformatics.</article-title>
                    <source>
						
                        <italic toggle="yes">Hum Genet.</italic>
					</source>
                    <year>2010</year>;<volume>128</volume>(<issue>2</issue>):<fpage>179</fpage>&#x2013;<lpage>194</lpage>.
                    <pub-id pub-id-type="pmid">20512354</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s00439-010-0837-0</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-15">
                <label>15</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Radivojac</surname>
                            <given-names>P</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Clark</surname>
                            <given-names>WT</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Oron</surname>
                            <given-names>TR</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>A large-scale evaluation of computational protein function prediction.</article-title>
                    <source>
						
                        <italic toggle="yes">Nat Methods.</italic>
					</source>
                    <year>2013</year>;<volume>10</volume>(<issue>3</issue>):<fpage>221</fpage>&#x2013;<lpage>7</lpage>.
                    <pub-id pub-id-type="pmid">23353650</pub-id>
                    <pub-id pub-id-type="doi">10.1038/nmeth.2340</pub-id>
                    <pub-id pub-id-type="pmcid">3584181</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-16">
                <label>16</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>L&#x00f3;pez</surname>
                            <given-names>G</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Rojas</surname>
                            <given-names>A</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tress</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Assessment of predictions submitted for the CASP7 function prediction category.</article-title>
                    <source>
						
                        <italic toggle="yes">Proteins.</italic>
					</source>
                    <year>2007</year>;<volume>69</volume>(<issue>Suppl 8</issue>):<fpage>165</fpage>&#x2013;<lpage>174</lpage>.
                    <pub-id pub-id-type="pmid">17654548</pub-id>
                    <pub-id pub-id-type="doi">10.1002/prot.21651</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-17">
                <label>17</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Pellegrini-Calace</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Soro</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tramontano</surname>
                            <given-names>A</given-names>
                        </name>
					</person-group>:
                    <article-title>Revisiting the prediction of protein function at CASP6.</article-title>
                    <source>
						
                        <italic toggle="yes">FEBS J.</italic>
					</source>
                    <year>2006</year>;<volume>273</volume>(<issue>13</issue>):<fpage>2977</fpage>&#x2013;<lpage>2983</lpage>.
                    <pub-id pub-id-type="pmid">16759228</pub-id>
                    <pub-id pub-id-type="doi">10.1111/j.1742-4658.2006.05309.x</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-18">
                <label>18</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Gillis</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Pavlidis</surname>
                            <given-names>P</given-names>
                        </name>
					</person-group>:
                    <article-title>Characterizing the state of the art in the computational assignment of gene function: lessons from the first critical assessment of functional annotation (CAFA).</article-title>
                    <source>
						
                        <italic toggle="yes">BMC Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>14</volume>(<issue>Suppl 3</issue>):<fpage>S15</fpage>.
                    <pub-id pub-id-type="pmid">23630983</pub-id>
                    <pub-id pub-id-type="doi">10.1186/1471-2105-14-S3-S15</pub-id>
                    <pub-id pub-id-type="pmcid">3633048</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-19">
                <label>19</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Stolovitzky</surname>
                            <given-names>G</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Monroe</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Califano</surname>
                            <given-names>A</given-names>
                        </name>
					</person-group>:
                    <article-title>Dialogue on reverse-engineering assessment and methods: the DREAM of high-throughput pathway inference.</article-title>
                    <source>
						
                        <italic toggle="yes">Ann N Y Acad Sci.</italic>
					</source>
                    <year>2007</year>;<volume>1115</volume>:<fpage>1</fpage>&#x2013;<lpage>22</lpage>.
                    <pub-id pub-id-type="pmid">17925349</pub-id>
                    <pub-id pub-id-type="doi">10.1196/annals.1407.021</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-20">
                <label>20</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Venet</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Dumont</surname>
                            <given-names>JE</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Detours</surname>
                            <given-names>V</given-names>
                        </name>
					</person-group>:
                    <article-title>Most random gene expression signatures are significantly associated with breast cancer outcome.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS Comput Biol.</italic>
					</source>
                    <year>2011</year>;<volume>7</volume>(<issue>10</issue>):<fpage>e1002240</fpage>.
                    <pub-id pub-id-type="pmid">22028643</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002240</pub-id>
                    <pub-id pub-id-type="pmcid">3197658</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-21">
                <label>21</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Boulesteix</surname>
                            <given-names>AL</given-names>
                        </name>
					</person-group>:
                    <article-title>On representative and illustrative comparisons with real data in bioinformatics: response to the letter to the editor by Smith 
                        <italic toggle="yes">et al.</italic>
                    </article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>29</volume>(<issue>20</issue>):<fpage>2664</fpage>&#x2013;<lpage>2666</lpage>.
                    <pub-id pub-id-type="pmid">23929033</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/btt458</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-22">
                <label>22</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Hoehndorf</surname>
                            <given-names>R</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Schofield</surname>
                            <given-names>PN</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gkoutos</surname>
                            <given-names>GV</given-names>
                        </name>
					</person-group>:
                    <article-title>An integrative, translational approach to understanding rare and orphan genetically based diseases.</article-title>
                    <source>
						
                        <italic toggle="yes">Interface Focus.</italic>
					</source>
                    <year>2013</year>;<volume>3</volume>(<issue>2</issue>):<fpage>20120055</fpage>.
                    <pub-id pub-id-type="pmid">23853703</pub-id>
                    <pub-id pub-id-type="doi">10.1098/rsfs.2012.0055</pub-id>
                    <pub-id pub-id-type="pmcid">3638468</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-23">
                <label>23</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Youngs</surname>
                            <given-names>N</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Penfold-Brown</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Drew</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Parametric Bayesian priors and better choice of negative examples improve protein function prediction.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>29</volume>(<issue>9</issue>):<fpage>1190</fpage>&#x2013;<lpage>8</lpage>.
                    <pub-id pub-id-type="pmid">23511543</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/btt110</pub-id>
                    <pub-id pub-id-type="pmcid">3634187</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-24">
                <label>24</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>King</surname>
                            <given-names>OD</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Lee</surname>
                            <given-names>JC</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Dudley</surname>
                            <given-names>AM</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Predicting phenotype from patterns of annotation.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2003</year>;<volume>19</volume>(<issue>Suppl 1</issue>):<fpage>i183</fpage>&#x2013;<lpage>189</lpage>.
                    <pub-id pub-id-type="pmid">12855456</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/btg1024</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-25">
                <label>25</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Lee</surname>
                            <given-names>I</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Li</surname>
                            <given-names>Z</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Marcotte</surname>
                            <given-names>EM</given-names>
                        </name>
					</person-group>:
                    <article-title>An improved, bias-reduced probabilistic functional gene network of baker&#x2019;s yeast, Saccharomyces cerevisiae.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2007</year>;<volume>2</volume>(<issue>10</issue>):<fpage>e988</fpage>.
                    <pub-id pub-id-type="pmid">17912365</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0000988</pub-id>
                    <pub-id pub-id-type="pmcid">1991590</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-26">
                <label>26</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Dolinski</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Botstein</surname>
                            <given-names>D</given-names>
                        </name>
					</person-group>:
                    <article-title>Automating the construction of gene ontologies.</article-title>
                    <source>
						
                        <italic toggle="yes">Nat Biotechnol.</italic>
					</source>
                    <year>2013</year>;<volume>31</volume>(<issue>1</issue>):<fpage>34</fpage>&#x2013;<lpage>35</lpage>.
                    <pub-id pub-id-type="pmid">23302932</pub-id>
                    <pub-id pub-id-type="doi">10.1038/nbt.2476</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-27">
                <label>27</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Dutkowski</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Kramer</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Surma</surname>
                            <given-names>MA</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>A gene ontology inferred from molecular networks.</article-title>
                    <source>
						
                        <italic toggle="yes">Nat Biotechnol.</italic>
					</source>
                    <year>2013</year>;<volume>31</volume>(<issue>1</issue>):<fpage>38</fpage>&#x2013;<lpage>45</lpage>.
                    <pub-id pub-id-type="pmid">23242164</pub-id>
                    <pub-id pub-id-type="doi">10.1038/nbt.2463</pub-id>
                    <pub-id pub-id-type="pmcid">3654867</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-28">
                <label>28</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Lage</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Hansen</surname>
                            <given-names>NT</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Karlberg</surname>
                            <given-names>EO</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>A large-scale analysis of tissue-specific pathology and gene expression of human disease genes and complexes.</article-title>
                    <source>
						
                        <italic toggle="yes">Proc Natl Acad Sci U S A.</italic>
					</source>
                    <year>2008</year>;<volume>105</volume>(<issue>52</issue>):<fpage>20870</fpage>&#x2013;<lpage>20875</lpage>.
                    <pub-id pub-id-type="pmid">19104045</pub-id>
                    <pub-id pub-id-type="doi">10.1073/pnas.0810772105</pub-id>
                    <pub-id pub-id-type="pmcid">2606902</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-29">
                <label>29</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Gillis</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Pavlidis</surname>
                            <given-names>P</given-names>
                        </name>
					</person-group>:
                    <article-title>Assessing identity, redundancy and confounds in Gene Ontology annotations over time.</article-title>
                    <source>
						
                        <italic toggle="yes">Bioinformatics.</italic>
					</source>
                    <year>2013</year>;<volume>29</volume>(<issue>4</issue>):<fpage>476</fpage>&#x2013;<lpage>482</lpage>.
                    <pub-id pub-id-type="pmid">23297035</pub-id>
                    <pub-id pub-id-type="doi">10.1093/bioinformatics/bts727</pub-id>
                    <pub-id pub-id-type="pmcid">3570208</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-30">
                <label>30</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Lipchina</surname>
                            <given-names>I</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Elkabetz</surname>
                            <given-names>Y</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Hafner</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Genome-wide identification of microRNA targets in human ES cells reveals a role for miR-302 in modulating BMP response.</article-title>
                    <source>
						
                        <italic toggle="yes">Genes Dev.</italic>
					</source>
                    <year>2011</year>;<volume>25</volume>(<issue>20</issue>):<fpage>2173</fpage>&#x2013;<lpage>2186</lpage>.
                    <pub-id pub-id-type="pmid">22012620</pub-id>
                    <pub-id pub-id-type="doi">10.1101/gad.17221311</pub-id>
                    <pub-id pub-id-type="pmcid">3205587</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-31">
                <label>31</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Mulvey</surname>
                            <given-names>CM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Tudzarova</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Crawford</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Subcellular proteomics reveals a role for nucleo-cytoplasmic trafficking at the DNA replication origin activation checkpoint.</article-title>
                    <source>
						
                        <italic toggle="yes">J Proteome Res.</italic>
					</source>
                    <year>2013</year>;<volume>12</volume>(<issue>3</issue>):<fpage>1436</fpage>&#x2013;<lpage>1453</lpage>.
                    <pub-id pub-id-type="pmid">23320540</pub-id>
                    <pub-id pub-id-type="doi">10.1021/pr3010919</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-32">
                <label>32</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>O&#x2019;Roak</surname>
                            <given-names>BJ</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Deriziotis</surname>
                            <given-names>P</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Lee</surname>
                            <given-names>C</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Exome sequencing in sporadic autism spectrum disorders identifies severe de novo mutations.</article-title>
                    <source>
						
                        <italic toggle="yes">Nat Genet.</italic>
					</source>
                    <year>2011</year>;<volume>43</volume>(<issue>6</issue>):<fpage>585</fpage>&#x2013;<lpage>589</lpage>.
                    <pub-id pub-id-type="pmid">21572417</pub-id>
                    <pub-id pub-id-type="doi">10.1038/ng.835</pub-id>
                    <pub-id pub-id-type="pmcid">3115696</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-33">
                <label>33</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Sookoian</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Pirola</surname>
                            <given-names>CJ</given-names>
                        </name>
					</person-group>:
                    <article-title>Metabolic syndrome: from the genetics to the pathophysiology.</article-title>
                    <source>
						
                        <italic toggle="yes">Curr Hypertens Rep.</italic>
					</source>
                    <year>2011</year>;<volume>13</volume>(<issue>2</issue>):<fpage>149</fpage>&#x2013;<lpage>157</lpage>.
                    <pub-id pub-id-type="pmid">20957457</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s11906-010-0164-9</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-34">
                <label>34</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Veerappa</surname>
                            <given-names>AM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Vishweswaraiah</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Lingaiah</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Unravelling the complexity of human olfactory receptor repertoire by copy number analysis across population using high resolution arrays.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>(<issue>7</issue>):<fpage>e66843</fpage>.
                    <pub-id pub-id-type="pmid">23843967</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0066843</pub-id>
                    <pub-id pub-id-type="pmcid">3700933</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-35">
                <label>35</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Kumimoto</surname>
                            <given-names>RW</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Siriwardana</surname>
                            <given-names>CL</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gayler</surname>
                            <given-names>KK</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>NUCLEAR FACTORY transcription factors have both opposing and additive roles in ABA-mediated seed germination.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>(<issue>3</issue>):<fpage>e59481</fpage>.
                    <pub-id pub-id-type="pmid">23527203</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0059481</pub-id>
                    <pub-id pub-id-type="pmcid">3602376</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-36">
                <label>36</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Franceschini</surname>
                            <given-names>A</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Szklarczyk</surname>
                            <given-names>D</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Frankild</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>STRING v9.1: protein-protein interaction networks, with increased coverage and integration.</article-title>
                    <source>
						
                        <italic toggle="yes">Nucleic Acids Res.</italic>
					</source>
                    <year>2013</year>;<volume>41</volume>(<issue>Database issue</issue>):<fpage>D808</fpage>&#x2013;<lpage>D815</lpage>.
                    <pub-id pub-id-type="pmid">23203871</pub-id>
                    <pub-id pub-id-type="doi">10.1093/nar/gks1094</pub-id>
                    <pub-id pub-id-type="pmcid">3531103</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-37">
                <label>37</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Lee</surname>
                            <given-names>I</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Blom</surname>
                            <given-names>UM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Wang</surname>
                            <given-names>PI</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Prioritizing candidate disease genes by network-based boosting of genome-wide association data.</article-title>
                    <source>
						
                        <italic toggle="yes">Genome Res.</italic>
					</source>
                    <year>2011</year>;<volume>21</volume>(<issue>7</issue>):<fpage>1109</fpage>&#x2013;<lpage>1121</lpage>.
                    <pub-id pub-id-type="pmid">21536720</pub-id>
                    <pub-id pub-id-type="doi">10.1101/gr.118992.110</pub-id>
                    <pub-id pub-id-type="pmcid">3129253</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-38">
                <label>38</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Tacutu</surname>
                            <given-names>R</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Shore</surname>
                            <given-names>DE</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Budovsky</surname>
                            <given-names>A</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Prediction of C. elegans longevity genes by human and worm longevity networks.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2012</year>;<volume>7</volume>(<issue>10</issue>):<fpage>e48282</fpage>.
                    <pub-id pub-id-type="pmid">23144747</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0048282</pub-id>
                    <pub-id pub-id-type="pmcid">3483217</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-39">
                <label>39</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Putnam</surname>
                            <given-names>CD</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Allen-Soltero</surname>
                            <given-names>SR</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Martinez</surname>
                            <given-names>SL</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Bioinformatic identification of genes suppressing genome instability.</article-title>
                    <source>
						
                        <italic toggle="yes">Proc Natl Acad Sci U S A.</italic>
					</source>
                    <year>2012</year>;<volume>109</volume>(<issue>47</issue>):<fpage>E3251</fpage>&#x2013;<lpage>E3259</lpage>.
                    <pub-id pub-id-type="pmid">23129647</pub-id>
                    <pub-id pub-id-type="doi">10.1073/pnas.1216733109</pub-id>
                    <pub-id pub-id-type="pmcid">3511103</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-40">
                <label>40</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Borra</surname>
                            <given-names>VM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Waterval</surname>
                            <given-names>JJ</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Stokroos</surname>
                            <given-names>RJ</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Localization of the gene for hyperostosis cranialis interna to chromosome 8p21 with analysis of three candidate genes.</article-title>
                    <source>
						
                        <italic toggle="yes">Calcif Tissue Int.</italic>
					</source>
                    <year>2013</year>;<volume>93</volume>(<issue>1</issue>):<fpage>93</fpage>&#x2013;<lpage>100</lpage>.
                    <pub-id pub-id-type="pmid">23640157</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s00223-013-9732-8</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-41">
                <label>41</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Breckpot</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Thienpont</surname>
                            <given-names>B</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Bauters</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Congenital heart defects in a novel recurrent 22q11.2 deletion harboring the genes CRKL and MAPK1.</article-title>
                    <source>
						
                        <italic toggle="yes">Am J Med Genet A.</italic>
					</source>
                    <year>2012</year>;<volume>158A</volume>(<issue>3</issue>):<fpage>574</fpage>&#x2013;<lpage>580</lpage>.
                    <pub-id pub-id-type="pmid">22318985</pub-id>
                    <pub-id pub-id-type="doi">10.1002/ajmg.a.35217</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-42">
                <label>42</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Chabchoub</surname>
                            <given-names>E</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Cogulu</surname>
                            <given-names>O</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Durmaz</surname>
                            <given-names>B</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Oculocerebral hypopigmentation syndrome maps to chromosome 3q27.1q29.</article-title>
                    <source>
						
                        <italic toggle="yes">Dermatology.</italic>
					</source>
                    <year>2011</year>;<volume>223</volume>(<issue>4</issue>):<fpage>306</fpage>&#x2013;<lpage>310</lpage>.
                    <pub-id pub-id-type="pmid">22327602</pub-id>
                    <pub-id pub-id-type="doi">10.1159/000335609</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-43">
                <label>43</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Chang</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Zhang</surname>
                            <given-names>W</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gao</surname>
                            <given-names>L</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Prioritization of candidate genes for attention deficit hyperactivity disorder by computational analysis of multiple data sources.</article-title>
                    <source>
						
                        <italic toggle="yes">Protein Cell.</italic>
					</source>
                    <year>2012</year>;<volume>3</volume>(<issue>7</issue>):<fpage>526</fpage>&#x2013;<lpage>534</lpage>.
                    <pub-id pub-id-type="pmid">22773342</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s13238-012-2931-7</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-44">
                <label>44</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Hitz</surname>
                            <given-names>MP</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Lemieux-Perreault</surname>
                            <given-names>LP</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Marshall</surname>
                            <given-names>C</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Rare copy number variants contribute to congenital left-sided heart disease.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS Genet.</italic>
					</source>
                    <year>2012</year>;<volume>8</volume>(<issue>9</issue>):<fpage>e1002903</fpage>.
                    <pub-id pub-id-type="pmid">22969434</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pgen.1002903</pub-id>
                    <pub-id pub-id-type="pmcid">3435243</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-45">
                <label>45</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>LopezJimenez</surname>
                            <given-names>N</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Gerber</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Popovici</surname>
                            <given-names>V</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Examination of FGFRL1 as a candidate gene for diaphragmatic defects at chromosome 4p16.3 shows that Fgfrl1 null mice have reduced expression of Tpm3, sarcomere genes and Lrtm1 in the diaphragm.</article-title>
                    <source>
						
                        <italic toggle="yes">Hum Genet.</italic>
					</source>
                    <year>2010</year>;<volume>127</volume>(<issue>3</issue>):<fpage>325</fpage>&#x2013;<lpage>336</lpage>.
                    <pub-id pub-id-type="pmid">20024584</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s00439-009-0777-8</pub-id>
                    <pub-id pub-id-type="pmcid">2893560</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-46">
                <label>46</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Melchionda</surname>
                            <given-names>L</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Fang</surname>
                            <given-names>M</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Wang</surname>
                            <given-names>H</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Adult-onset alexander disease, associated with a mutation in an alternative GFAP transcript, may be phenotypically modulated by a non-neutral HDAC6 variant.</article-title>
                    <source>
						
                        <italic toggle="yes">Orphanet J Rare Dis.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>:<fpage>66</fpage>.
                    <pub-id pub-id-type="pmid">23634874</pub-id>
                    <pub-id pub-id-type="doi">10.1186/1750-1172-8-66</pub-id>
                    <pub-id pub-id-type="pmcid">3654953</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-47">
                <label>47</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Wang</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Qian</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Hoeksema</surname>
                            <given-names>MD</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Integrative genomics analysis identifies candidate drivers at 3q26-29 amplicon in squamous cell carcinoma of the lung.</article-title>
                    <source>
						
                        <italic toggle="yes">Clin Cancer Res.</italic>
					</source>
                    <year>2013</year>;<volume>19</volume>(<issue>20</issue>):<fpage>5580</fpage>&#x2013;<lpage>5590</lpage>.
                    <pub-id pub-id-type="pmid">23908357</pub-id>
                    <pub-id pub-id-type="doi">10.1158/1078-0432.CCR-13-0594</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-48">
                <label>48</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Zhu</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Cui</surname>
                            <given-names>L</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Wang</surname>
                            <given-names>W</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Whole exome sequencing identifies mutation of EDNRA involved in ACTH-independent macronodular adrenal hyperplasia.</article-title>
                    <source>
						
                        <italic toggle="yes">Fam Cancer.</italic>
					</source>
                    <year>2013</year>.
                    <pub-id pub-id-type="pmid">23754170</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s10689-013-9642-y</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-49">
                <label>49</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Ho</surname>
                            <given-names>DW</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Yap</surname>
                            <given-names>MK</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Ng</surname>
                            <given-names>PW</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Association of high myopia with crystallin beta A4 (CRYBA4) gene polymorphisms in the linkage-identified MYP6 locus.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2012</year>;<volume>7</volume>(<issue>6</issue>):<fpage>e40238</fpage>.
                    <pub-id pub-id-type="pmid">22792142</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0040238</pub-id>
                    <pub-id pub-id-type="pmcid">3389832</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-50">
                <label>50</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Hussain</surname>
                            <given-names>MS</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Baig</surname>
                            <given-names>SM</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Neumann</surname>
                            <given-names>S</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>A truncating mutation of CEP135 causes primary microcephaly and disturbed centrosomal function.</article-title>
                    <source>
						
                        <italic toggle="yes">Am J Hum Genet.</italic>
					</source>
                    <year>2012</year>;<volume>90</volume>(<issue>5</issue>):<fpage>871</fpage>&#x2013;<lpage>878</lpage>.
                    <pub-id pub-id-type="pmid">22521416</pub-id>
                    <pub-id pub-id-type="doi">10.1016/j.ajhg.2012.03.016</pub-id>
                    <pub-id pub-id-type="pmcid">3376485</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-51">
                <label>51</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Thiel</surname>
                            <given-names>C</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Kessler</surname>
                            <given-names>K</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Giessl</surname>
                            <given-names>A</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>NEK1 mutations cause short-rib polydactyly syndrome type majewski.</article-title>
                    <source>
						
                        <italic toggle="yes">Am J Hum Genet.</italic>
					</source>
                    <year>2011</year>;<volume>88</volume>(<issue>1</issue>):<fpage>106</fpage>&#x2013;<lpage>114</lpage>.
                    <pub-id pub-id-type="pmid">21211617</pub-id>
                    <pub-id pub-id-type="doi">10.1016/j.ajhg.2010.12.004</pub-id>
                    <pub-id pub-id-type="pmcid">3014367</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-52">
                <label>52</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Yu</surname>
                            <given-names>L</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Wynn</surname>
                            <given-names>J</given-names>
                        </name>
						
                        <name name-style="western">
                            <surname>Cheung</surname>
                            <given-names>YH</given-names>
                        </name>
						
                        <etal/>
					</person-group>:
                    <article-title>Variants in GATA4 are a rare cause of familial and sporadic congenital diaphragmatic hernia.</article-title>
                    <source>
						
                        <italic toggle="yes">Hum Genet.</italic>
					</source>
                    <year>2013</year>;<volume>132</volume>(<issue>3</issue>):<fpage>285</fpage>&#x2013;<lpage>292</lpage>.
                    <pub-id pub-id-type="pmid">23138528</pub-id>
                    <pub-id pub-id-type="doi">10.1007/s00439-012-1249-0</pub-id>
                    <pub-id pub-id-type="pmcid">3570587</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-53">
                <label>53</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Michailidis</surname>
                            <given-names>G</given-names>
                        </name>
					</person-group>:
                    <article-title>Statistical challenges in biological networks.</article-title>
                    <source>
						
                        <italic toggle="yes">J Comput Graph Stat.</italic>
					</source>
                    <year>2012</year>;<volume>21</volume>(<issue>4</issue>):<fpage>840</fpage>&#x2013;<lpage>855</lpage>.
                    <pub-id pub-id-type="doi">10.1080/10618600.2012.738614</pub-id>
                </mixed-citation>
            </ref>
            <ref id="ref-54">
                <label>54</label>
                <mixed-citation publication-type="journal">
                    <person-group person-group-type="author">
						
                        <name name-style="western">
                            <surname>Vey</surname>
                            <given-names>G</given-names>
                        </name>
					</person-group>:
                    <article-title>Metagenomic guilt by association: an operonic perspective.</article-title>
                    <source>
						
                        <italic toggle="yes">PLoS One.</italic>
					</source>
                    <year>2013</year>;<volume>8</volume>(<issue>8</issue>):<fpage>e71484</fpage>.
                    <pub-id pub-id-type="pmid">23940763</pub-id>
                    <pub-id pub-id-type="doi">10.1371/journal.pone.0071484</pub-id>
                    <pub-id pub-id-type="pmcid">3735515</pub-id>
                </mixed-citation>
            </ref>
        </ref-list>
    </back>
    <sub-article article-type="reviewer-report" id="report3158">
        <front-stub>
            <article-id pub-id-type="doi">10.5256/f1000research.2457.r3158</article-id>
            <title-group>
                <article-title>Reviewer response for version 1</article-title>
            </title-group>
            <contrib-group>
                <contrib contrib-type="author">
                    <name>
                        <surname>Anantharaman</surname>
                        <given-names>Vivek</given-names>
                    </name>
                    <xref ref-type="aff" rid="r3158a1">1</xref>
                    <role>Referee</role>
                </contrib>
                <aff id="r3158a1">
                    <label>1</label>National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD, USA</aff>
            </contrib-group>
            <author-notes>
                <fn fn-type="conflict">
                    <p>
                        <bold>Competing interests: </bold>No competing interests were disclosed.</p>
                </fn>
            </author-notes>
            <pub-date pub-type="epub">
                <day>5</day>
                <month>2</month>
                <year>2014</year>
            </pub-date>
            <permissions>
                <copyright-statement>Copyright: &#x00a9; 2014 Anantharaman V</copyright-statement>
                <copyright-year>2014</copyright-year>
                <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
                    <license-p>This is an open access peer review report distributed under the terms of the Creative Commons Attribution Licence, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
                </license>
            </permissions>
            <related-article ext-link-type="doi" id="relatedArticleReport3158" related-article-type="peer-reviewed-article" xlink:href="10.12688/f1000research.2-230.v1"/>
            <custom-meta-group>
                <custom-meta>
                    <meta-name>recommendation</meta-name>
                    <meta-value>approve</meta-value>
                </custom-meta>
            </custom-meta-group>
        </front-stub>
        <body>
            <p>The authors have discussed the pitfalls of automated GBA and ways to improve functional prediction. Automated methods for GBA are only as good as ontologies and curated reference datasets. Ontologies like GO suffer from poor quality annotation being propagated throughout their data that result in &#x201c;Garbage in &#x2013; Garbage out&#x201d; phenomenon. A generic functional prediction is the best one can expect from existing automated methods.</p>
            <p>From my experience, I have found that accurate functional prediction requires a mix of local sequence similarity and sequence profile searches, proper sequence analysis with study of sequence and phyletic conservation, structural analysis, network studies of various data points, and correlation with experimental data, all done with a heavy dose of manual tuning. None of the automated methods of GBA give consistent accurate prediction, without manual intervention.</p>
            <p>The review laid out by the authors is a good analysis of the challenges and limitation of gene function prediction. One area that the authors do not explicitly discuss is the great difference between eukaryotes and prokaryotes. GBA is currently far more effective in the latter due to operons, which are not available in eukaryotes.</p>
            <p>Reviewer Expertise:</p>
            <p>NA</p>
            <p>I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.</p>
        </body>
    </sub-article>
    <sub-article article-type="reviewer-report" id="report2276">
        <front-stub>
            <article-id pub-id-type="doi">10.5256/f1000research.2457.r2276</article-id>
            <title-group>
                <article-title>Reviewer response for version 1</article-title>
            </title-group>
            <contrib-group>
                <contrib contrib-type="author">
                    <name>
                        <surname>Toppo</surname>
                        <given-names>Stefano</given-names>
                    </name>
                    <xref ref-type="aff" rid="r2276a1">1</xref>
                    <role>Referee</role>
                </contrib>
                <aff id="r2276a1">
                    <label>1</label>Department of Molecular Medicine, Universit&#x00e0; degli studi di Padova, Padova, Italy</aff>
            </contrib-group>
            <author-notes>
                <fn fn-type="conflict">
                    <p>
                        <bold>Competing interests: </bold>No competing interests were disclosed.</p>
                </fn>
            </author-notes>
            <pub-date pub-type="epub">
                <day>8</day>
                <month>11</month>
                <year>2013</year>
            </pub-date>
            <permissions>
                <copyright-statement>Copyright: &#x00a9; 2013 Toppo S</copyright-statement>
                <copyright-year>2013</copyright-year>
                <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
                    <license-p>This is an open access peer review report distributed under the terms of the Creative Commons Attribution Licence, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
                </license>
            </permissions>
            <related-article ext-link-type="doi" id="relatedArticleReport2276" related-article-type="peer-reviewed-article" xlink:href="10.12688/f1000research.2-230.v1"/>
            <custom-meta-group>
                <custom-meta>
                    <meta-name>recommendation</meta-name>
                    <meta-value>approve</meta-value>
                </custom-meta>
            </custom-meta-group>
        </front-stub>
        <body>
            <p>This opinion article deals with the long standing issue of protein function prediction in its broader sense. The authors express an interesting and most of the time shareable point of view about the negative impact of gene multifunctionality that influences gene network-based guilt-by-association studies.</p>
            <p>The paper is really well written and organized in sections that focus on different aspects of function prediction and its pitfalls. Nonetheless, there are some minor points that I would suggest mitigating, as they sound too harsh and are, as far as I&#x2019;m concerned, partly incorrect.</p>
            <p>In the &#x201c;
                <italic>Finding better algorithms</italic>&#x201d; section, CAFA is mentioned and commented on but I would like to pinpoint some aspects about how this is done and some related issues (below). Even the glorious series of CASP experiments, that the authors have mentioned, suffered a lot in their first editions but what is more important is that both assessors and participants are aware of this and that improvements are planned, as far as I know.</p>
            <p>The prediction results from the CAFA experiment could perhaps be framed around some different points of view:</p>
            <p>
                <bold>1:</bold>
            </p>
            <p>Looking at the F-measure results for top performing methods there is little to be happy about. There are the following additional issues coming out from CAFA: are we really sure that some of the predicted functions are not correct? Is this rather an effect caused by the possible incompleteness of some experimental data? In other words, can anyone assert firmly that there is nothing else to discover about the function of a protein? I would definitively say NO. There is more than meets the eye and besides, the benchmark is incomplete by definition as it will never complete in the future either, no matter what information is added. Not only that, but even novel experimental evidence can turn false positive predictions of protein function into a true positive. On the flip side, true positive predictions can equally be refuted by fresh experimental data and consequently turn into a false positive.</p>
            <p>
                <bold>2:</bold>
            </p>
            <p>Some functions of the CAFA experiment were extremely difficult to predict and hard to &#x201c;guess&#x201d; in any way, both from a simple sequence similarity approach or other more sophisticated techniques based on machine learning. In summary, some CAFA targets were not so easy to predict.</p>
            <p>
                <bold>3:</bold>
            </p>
            <p>Additionally with CAFA, much marginally informative experimental data was collected for many targets that mainly derived from PPI experiments. The term under indictment is &#x201c;protein binding&#x201d; which is heavily present, for example, in the GOA database of annotated proteins. The assessors&#x2019; decision to discard &#x201c;protein binding&#x201d; in the final evaluation was, consequently, correct. The good performance of na&#x00ef;ve and BLAST methods (when &#x201c;protein binding&#x201d; is considered in the assessment) depends on the high occurrence (multifunctional?) of one function in the database and its prevalence over the others so that it is very easy &#x201c;to predict&#x201d;. In this sense, it is not exactly correct to assert that BLAST and na&#x00ef;ve are (almost) the best performing tools because many tools participating to CAFA, whenever possible, tried to provide more informative annotations in place of the less informative &#x201c;protein binding&#x201d;. Taking this into account, &#x201c;protein binding&#x201d; would have been inappropriate to use in the final evaluation because it would have rewarded BLAST and na&#x00ef;ve methods artificially but penalized others.</p>
            <p>In contrast, the main issue is that databases contain biased annotations and a few scarcely informative terms dominate the scene. In this respect, I totally agree with the authors that multifunctionality poses serious problems to function prediction algorithms. So how might one mitigate the effect of multifunctional and scarcely informative annotations?&#x00a0;</p>
            <p>Perhaps CAFA will need to settle in the next editions and the contribution to this process of renewal should be constructive and proactive rather than purely critical.</p>
            <p>
                <bold>4:</bold>
            </p>
            <p>The authors recognize that successful stories may be limited, simple and hand-tuned. Reverse engineering results is demanding and I agree with the authors, but I would note that the excess in the analysis, as suggested in the when/how methods perform section, could lead to an overestimate/underestimate of the behavior of the tools and miss their general action. As a matter of fact, biology is made of more exceptions than rules and tools are designed to follow only the rules. Can the authors suggest some possible ways in which a tentative solution can be set up that could be discussed and adopted in critical assessments of function prediction tools?</p>
            <p>
                <bold>5:</bold>
            </p>
            <p>To me, the distinction in this paper between between GO and protein annotations using GO are not clear enough. I would, for instance, rephrase the following:&#x00a0;</p>
            <p>&#x201c;&#x2026;. 
                <italic>We also suggested that part of the problem is the reliance by computational biologists on gold standard annotations such as the Gene Ontology &#x2026;.&#x201d;</italic>
            </p>
            <p>To something like:</p>
            <p>&#x201c;&#x2026;. We also suggested that part of the problem is the reliance by computational biologists on gold standard annotations such as the Gene Ontology Annotation database (GOA) &#x2026;.&#x201d;</p>
            <p>I know the authors know the difference between the two and for this reason I would recommend that they clarify this aspect and do not confound what GO and its countless instances are (one of them is GOA).</p>
            <p>As an obvious reminder, GO is an abstraction of the knowledge tentatively organized in a directed acyclic graph and is a controlled vocabulary intended as the rosetta stone of different interpretations and expressions of the same concepts. I strongly believe that GO is rigorous and we can trust it. On the contrary, GOA contains GO instances used to describe proteins. Using the metaphor of programming language, the GO term is the &#x201c;object&#x201d; and its use in GOA, or other databases containing GO annotated proteins, is the &#x201c;instance&#x201d; of that &#x201c;object&#x201d;. This is an important difference because the &#x201c;object&#x201d; is abstract and may be varied in a number of ways that can be right or wrong when thinking about protein annotation. It is not the GO term definition per se in the dock but rather its utilization as a descriptor of protein function stored in public databases. The authors already published a paper on GO based annotations and their distribution in GOA over time. In other words, one can rely on the GO descriptions and their positions in the graph (though GO is continuously revisited, it is rather stable) but must pay attention to the proteins annotated with GO terms because they can be inappropriate and change over time, as already evaluated by the authors in a previous work (indeed, GO annotations in GOA change frequently).</p>
            <p>It would be very interesting to know what the authors think about the latter phenomenon, i.e. the updating of old annotations and their syncing with novel and more precise GO ontologies.</p>
            <p>Most of the issues, fully and carefully described by the authors, may be due more to this aspect than others. Generic annotations may have been used at the beginning of the story when GO was still incomplete and with poor coverage of biological knowledge. Thinking about &#x201c;Inferred from Electronic Annotations&#x201d; (IEA), these GO terms may have created the multifunctional phenomenon as they had time to spread and have consequently become both pervasive and difficult to eradicate or update.</p>
            <p>Reviewer Expertise:</p>
            <p>NA</p>
            <p>I confirm that I have read this submission and believe that I have an appropriate level of expertise to confirm that it is of an acceptable scientific standard.</p>
        </body>
    </sub-article>
</article>
