Showing posts with label Bioinformatics. Show all posts
Showing posts with label Bioinformatics. Show all posts

Monday, August 17, 2015

PhD thesis lengths

Bioinformatics lies at the nexus of the biological sciences and the computational sciences. Therefore it is sometimes worth comparing these two disciplines.

Marcus Beck at the R is My Friend blog has looked at doctoral dissertation lengths via the digital archives at the University of Minnesota. His data are shown in this box plot. You can search through it for your own favorite discipline (click on the image to make it larger).


He also has several other graphical views in his blog post, including data on masters theses.

Wednesday, November 12, 2014

Archiving of phylogenetics data


The draft Minimum Information about a Phylogenetic Analysis standard (Leebens-Mack et al. 2006) suggests that all relevant information about each and every published phylogenetics analysis should be archived, so that it can be scrutinized by later researchers, either for validation or for re-use. The issues here are both preservation of the information (data and analysis protocols) and open access to it.

In this blog we have already pointed out that there has been criticism of the bioinformatics part of this archiving, where there have been repeated claims that many computer programs are poorly maintained (Poor bioinformatics?) as well as poorly archived (Archiving of bioinformatics software).

Anyone who has ever tried to get data out of a biologist will know that the data-related part of the standard is no better. My own success rate, at requesting data from all areas of biology not just phylogenetics, is less than 20% over the past 25 years. The responses have been, in order: (i) no response (>50%), (ii) "a student / postdoc / colleague has the data not me", and (iii) "I have moved recently and don't know where the data are". My most recent attempt, to get the data from Collard et al. (2006), was ultimately unsuccessful even after several attempts.


For phylogenetics, this situation has recently been quantified and analyzed by Magee et al. (2014). They tried to collect phylogenetic data (comprising nucleotide sequence alignment and tree files) from 217 published studies. Of these, 54 (25%) had at least some part of the data (alignment or tree) archived in an online repository, and 91 (42%) were obtained by direct solicitation, but in 72 (33%) of cases nothing could be obtained even after three requests. Overall, complete datasets (both tree and alignment) were available for only 40% of the studies.

The authors note that the data were more likely to be deposited in online archives and/ or shared upon request when the publishing journal has a strong data-sharing policy. Furthermore, there has been a positive impact of recent policy initiatives and infrastructural changes involving data repositories. The TreeBASE phylogenetic-data repository has existed for more than 20 years, but its use has been sporadic. However, the recent establishment of the Joint Data Archiving Policy by a consortium of journals, which requires the submission of data to online archives as a condition of publication, and the concomitant establishment of the Dryad repository for evolutionary and ecological data, has seen a surge in the archiving of data.

So, all in all, things have been no better on the bio side than the informatics side of bioinformatics.

Stoltzfus et al. (2012) have identified a number of possible barriers to successful data archiving, including lack of awareness of options and policies, perception that benefits do not justify burden, and an active desire to restrict data access. Importantly, there are also a number of practical issues even for those people who do wish to archive their data:
  • inconvenience of gathering complete data and metadata
  • inconvenience of format conversions needed for archiving
  • frustration when some data don't fit the archive's data model
  • poor and undocumented archive submission interfaces.
For the readers of this blog, issue three is possibly the most important one — all current repositories are based on a tree model for phylogenetics, and therefore network phylogenies are frustrating to deal with.

In order to improve the overall situation, there are explicit suggestions from Cranston et al. (2014) for best practices when archiving. They have ten simple guidelines that, if followed, will result in you providing open access to your data and analyses, even if the publishing journal does not force you to do it.

Footnote: I have been reminded that archiving data in PDF format is inappropriate. Trying to extract text (such as a dataset) from a PDF file can be difficult, because there is no standard format for storing the text. Consequently, different PDF readers will extract the text in different ways, and it is possible that in all cases the output will need extensive manual re-formatting, in order to recover the original text formatting that went into the PDF file. In my experience, Google Chrome may do the least-worst job.

References

Collard M, Shennan SJ, Tehrani JJ (2006) Branching, blending, and the evolution of cultural similarities and differences among human populations. Evolution and Human Behavior 27: 169-184.

Cranston K, Harmon LJ, O'Leary MA, Lisle C (2014) Best practices for data sharing in phylogenetic research. PLoS Currents Jun 19;6.

Leebens-Mack J, Vision T, Brenner E, Bowers JE, Cannon S, Clement MJ, Cunningham CW, dePamphilis C, deSalle R, Doyle JJ, Eisen JA, Gu X, Harshman J, Jansen RK, Kellogg EA, Koonin EV, Mishler BD, Philippe H, Pires JC, Qiu YL, Rhee SY, Sjölander K, Soltis DE, Soltis PS, Stevenson DW, Wall K, Warnow T, Zmasek C (2006) Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA). OMICS 10: 231-237.

Magee AF, May MR, Moore BR (2014) The dawn of open access to phylogenetic data. PLoS One 9: e110268.

Stoltzfus A, O'Meara B, Whitacre J, Mounce R, Gillespie EL, Kumar S, Rosauer DF, Vos RA (2012) Sharing and re-use of phylogenetic trees (and associated data) to facilitate synthesis. BMC Research Notes 5: 574.

Thursday, November 6, 2014

Massive citations of bioinformatics in biology papers


For those of you who have missed it, the magazine Nature has recently looked at the 100 most highly cited science papers of all time (across all fields):
van Noorden R, Maher B, Nuzzo R (2014) The top 100 papers: Nature explores the most-cited research of all time. Nature 514: 550-553.
The list is dominated by biology papers, with biochemical laboratory techniques taking all of the top spots. However, it also worth noting that bioinformatics papers produce a very good showing, and so I have extracted 10 of them here.

If you have ever wondered what phylogenetic tree-building method is most used then it is at #20, while the most-used tree-building program is at #45 (having got there in only 7 years). You may also wonder why sequence alignment programs (#10 & #28 for Clustal; #12 & #14 for BLAST) do much better than tree-building programs (#45 for MEGA; #75 for GCG; #100 for MrBayes).

As for journals, the papers appeared in Nucleic Acids Research (4), Molecular Biology & Evolution (2), Bioinformatics (2), Journal of Molecular Biology (1) and Evolution (1). This list only partially matches their Journal Citation Reports current 5-Year Impact Factors: 8.378, 10.494, 6.968, 3.795 and 5.469, respectively.


Rank: 10 Citations: 40,289
Clustal W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
Thompson, J. D., Higgins, D. G. & Gibson, T. J
Nucleic Acids Res. 22, 4673–4680 (1994).

Rank: 12 Citations: 38,380
Basic local alignment search tool.
Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J.
J. Mol. Biol. 215, 403–410 (1990).

Rank: 14 Citations: 36,410
Gapped BLAST and PSI-BLAST: A new generation of protein database search programs.
Altschul, S. F. et al.
Nucleic Acids Res. 25, 3389–3402 (1997).

Rank: 20 Citations: 30,176
The neighbor-joining method: A new method for reconstructing phylogenetic trees.
Saitou, N. & Nei, M.
Mol. Biol. Evol. 4, 406–425 (1987).

Rank: 28 Citations: 24,098
The CLUSTAL_X Windows interface: Flexible strategies for multiple sequence alignment aided by quality analysis tools.
Thompson, J. D., Gibson, T. J., Plewniak, F., Jeanmougin, F. & Higgins, D. G.
Nucleic Acids Res. 25, 4876–4882 (1997).

Rank: 41 Citations: 21,373
Confidence limits on phylogenies: an approach using the bootstrap.
Felsenstein, J.
Evolution 39, 783–791 (1985).

Rank: 45 Citations: 18,286
MEGA4: Molecular Evolutionary Genetics Analysis (MEGA) software version 4.0.
Tamura, K., Dudley, J., Nei, M. & Kumar, S.
Mol. Biol. Evol. 24, 1596–1599 (2007).

Rank: 75 Citations: 14,226
A comprehensive set of sequence analysis programs for the VAX.
Devereux, J., Haeberli, P. & Smithies, O.
Nucleic Acids Res. 12, 387–395 (1984).

Rank: 76 Citations: 14,099
MODELTEST: Testing the model of DNA substitution.
Posada, D. & Crandall, K. A.
Bioinformatics 14, 817–818 (1998).

Rank: 100 Citations: 12,209
MrBayes 3: Bayesian phylogenetic inference under mixed models.
Ronquist, F. & Huelsenbeck, J. P.
Bioinformatics 19, 1572–1574 (2003).

Monday, August 18, 2014

Bioinformaticians' nightmares


These illustrations are from Alper Uzun's Biocomicals web site.






Bioinformaticians' dream



Bioinformaticians' reality


Wednesday, May 14, 2014

Non-model distances in phylogenetics


Bioinformatics is sometimes divorced from biology. This happens when the ideas involved are solely computational ones, and are not biologically motivated. For example, Otu et al. (2003), when proposing "a new sequence distance measure for phylogenetic tree construction" point out that "it is worth noting that our distance measures do not use any evolutionary model". They seem to see this a a Good Thing whereas biologists might not.

Pairwise distances are actually a good case in point, because there are many ways to measure the similarity / distance between any two objects, and most of them have nothing whatever to do with biology. One method that was popular about 10 years ago was using computerized compression.


This is based on the notion of complexity — all objects are complex to one degree or another, but the information they contain can often be reduced (or compressed) to a smaller amount without losing anything essential. That is, the original information can be exactly reproduced from the compressed version. This notion is applied to computer files, for example, in the idea of zipping a file — depending on the file's content it may be possible to zip the file to much a smaller size for storage or transmission.

This idea is straightforward to apply to similarity (Cilibrasi & Vitányi 2005). One simply has to compress the two objects separately as well as compress the combination of the two objects (ie. their union or concatenation). If the two objects are identical then the compression of their union will be the same size as the compression of each object individually (= a distance of 0), because all of the information in one of the objects is redundant. If they have nothing in common then the compression of their union will be the sum of the sizes of the compression of each object individually (= a distance of 1).

This idea has been applied in a number of areas where similarity / distance is used to group objects together (ie. clustering or classification), including computer viruses (Wehner 2008), music (Cilibrasi et al. 2004), languages (Benedetto et al. 2002a), and genomics (Chen et al. 2000, Li et al. 2001, Otu et al. 2003).

This approach has not always been received with equanimity. For example, the paper by Benedetto et al. (2002a) has received considerable commentary (Khmelev & Teahan 2003a; Benedetto et al. 2003; Khmelev & Teahan 2003b), (Goodman 2002; Benedetto et al. 2002b), (Wang 2009).

In biology, this idea has been mostly ignored. Chen et al. (2000) and Li et al. (2001) based their idea of information on Kolmogorov complexity, which is "the shortest program on a universal computer that outputs one of the objects when the input is the other object". This complexity cannot be measured directly, and so it is approximated by using a file compression program (eg. gzip). Clearly, the approximation can only be as good as the compression program and its algorithms.

Otu et al. (2003) based their idea of information on Lempel & Ziv complexity, which is "related to the number of steps required by a production process that builds either of the two objects". They claim an exact solution to the measurement of this complexity for gene sequences, and propose several slightly different distances based on this. Furthermore, the claim to:
show that the proposed distance measures, which are based on the relative complexity between sequences imply the evolutionary distance between organisms ... As the proposed approach does not depend on multiple alignments we test the validity of the approach in two ways: we use simulated data to show that the proposed distance measures can reasonably be represented by a tree. We also show the superiority of the proposed method on existing techniques using this simulated data. Secondly, we look at how well the results generated by the proposed method agree with existing phylogenies.
The use of generic compression similarity simply demonstrates that phylogeny does produce similarity to some extent. However, phylogeny is based on what we might call a form of "special similarity", which is similarity solely due to the possession of shared derived character states. Other components of similarity, such as the similarity produced by parallelisms, convergences and reversals, do not contribute to an estimate of a phylogeny. Indeed, they will often be positively misleading. A general concept of similarity is inadequate for reconstructing a phylogeny except under the simplest circumstances. This was empirically demonstrated, for example, in liguistics by the results of the Computer-Assisted Stemmatology Challenge, in which the compression method CompLearn came dead last in the primary ranking of the different phylogenetic methods (but did okay in the secondary ranking).

In short, only some aspects of complexity are relevant for phylogenetics, and only some of the information contained in a genome is relevant for measuring phylogenetic similarity. The focus on information as a whole ignores the "bio" in bioinformatics.

References

Benedetto D, Caglioti E, Loreto V (2002a) Language trees and zipping. Physical Review Letters 88: 048702. [arXiv:cond-mat/0108530, 2001]

Benedetto D, Caglioti E, Loreto V (2002b) On J. Goodman’s comment to "Language trees and zipping". arXiv:cond-mat/0203275

Benedetto D, Caglioti E, Loreto V (2003) A reply to the comment by Khmelev and Teahan. Physical Review Letters 90: 089804.

Chen X, Kwong S, Li M (2000) A compression algorithm for DNA sequences and its applications in genome comparison. Proceedings of the Fourth Annual International Conference on Computational Molecular Biology (RECOMB), pp. 107-117. ACM Press, Tokyo.

Cilibrasi R, Vitányi PMB (2005) Clustering by compression. IEEE Transactions on Information Theory 51: 1523-1545.

Cilibrasi R, Vitányi P, de Wolf R (2004) Algorithmic clustering of music based on string compression. Computer Music Journal 28(4): 49-67.

Goodman J (2002) Extended comment on "Language trees and zipping". arXiv:cond-mat/0202383

Khmelev DV, Teahan WJ (2003a) Comment on "Language trees and zipping". Physical Review Letters 90: 089803.

Khmelev DV, Teahan WJ (2003b) Comment on the reply of Benedetto et al. arXiv:cond-mat/0303261

Li M, Badger JH, Chen X, Kwong S, Kearney P, Zhang H (2001) An information-based sequence distance and its application to whole mitochondrial genome phylogeny. Bioinformatics 17: 149-154.

Otu HH, Sayood K (2003) A new sequence distance measure for phylogenetic tree construction. Bioinformatics 19: 2122-2130.

Wang X-L (2009) Comment on "Language trees and zipping". arXiv:0903.3669

Wehner S (2008) Analyzing worms and network traffic using compression. Journal of Computer Security 15: 303-320. (arXiv:cs/0504045v1, 2007)

Monday, January 13, 2014

Bioinformatics and inter-disciplinary work


Bioinformaticians are sometimes seen as multi-disciplinary workers (see the previous post on Results of some bioinformatics polls). If so, then the results of a recent study may be of interest:
Kevin M. Kniffin and Andrew S. Hanks (2013) Boundary spanning in academia: antecedents and near-term consequences of academic entrepreneurialism. Cornell Higher Education Research Institute Working Paper 158.
Kniffin and Hanks used data from the Survey of Earned Doctorates (conducted by the National Science Foundation), based on data from all people who earned PhDs in the U.S.A. between July 1 2009 and June 30 2010 (c. 43,000 people). Apparently, 14,000 people (32.5 %) reported that their doctoral work spanned academic boundaries

Two of their main findings are: (i) individuals who complete an interdisciplinary dissertation display short-term income risk, since they tend to earn nearly $1,700 less in the year after graduation; and (ii) the probability that non-citizens pursue interdisciplinary dissertation work is 4.7% higher when compared with U.S. citizens. Sadly, but not unexpectedly, women tend to earn less compared to men upon completion of the doctorate. Perhaps less expectedly, European American individuals also earn less in their first year after graduation than those in other racial groups.

For us, some of the more interesting data are:
Discipline
Agricultural & Life Sciences
Biological Sciences
Health Sciences
Computer Sciences & Mathematics
% of all Research Doctorates
2.3
17.6
4.4
7.0
% Interdisciplinary
44.5
41.1
29.9
22.7

In the regression models, adjusting for all other factors, the "influence of interdisciplinary research upon salary" was positive for Computer Sciences & Mathematics as well as for Health Sciences, but was negative for Biological Sciences. However, the "influence of interdisciplinary research upon employment as postdoctoral researcher" was negative for Computer Sciences & Mathematics as well as for Health Sciences, but was positive for Biological Sciences.

Make of this what you will.

Monday, December 9, 2013

Results of some bioinformatics polls


In 2008, Michael Barton conducted a Bioinformatics Career Survey. Since then, various groups have updated some of that information by conducting polls of their own. Below, I have included some of the more recent results, for your edification.

This first one comes from the Bioinformatics Organization, in response to the question: What is your undergraduate degree in? It is interesting to note that more bioinformaticians are biologists by training, rather than computational people.


The next one is actually an ongoing poll at BioCode's Notes, in response to the question: Which are the best programming languages for a bioinformatician? R is an interesting choice as the most useful language, given the more "traditional" use of Perl and Python.


That leads logically to another of the Bioinformatics Organization's questions: Which computer language are you most interested in learning (next) for bioinformatics R&D? I guess that if you already know R, then either Python or Perl is a useful thing to learn next.


Furthermore, the Bioinformatics Organization also asked: Which math / statistics language / application do you most frequently use? The choice of R here is more obvious, given that it is free, which most of the others are not. I wonder what the answer "none of the above" refers to.


Wednesday, November 20, 2013

Bioinformaticians look at bioinformatics


Bioinformatics as a term dates back to the 1970s, usually credited to Paulien Hogeweg, of the Bioinformatics group at Utrecht University, in The Netherlands, although it apparently did not make it into print until 1988 (Paulien Hogeweg. 1988. MIRROR beyond MIRROR, puddles of Life. In: Artificial Life, C. Langton, ed. Addison Wesley, pp. 297-315.).

In the 1990s the field expanded rapidly and became recognized as a discipline of its own, as a subset of computational science. However, Christos A. Ouzounis (2012. Rise and demise of bioinformatics? Promise and progress. PLoS Computational Biology 8: e1002487) has noted a distinct decrease in the use of the term itself, as shown by this graph.


Ouzounis recognizes three (admittedly artificial) periods in the history: Infancy (1996-2001), Adolescence (2002-2006) and Adulthood (2007-2011). Along the way, the practice of bioinformatics has received a lot of criticism. I have noted some of this before, in previous blog posts:
Poor bioinformatics?
Archiving of bioinformatics software

What is perhaps most important is that much of this criticism comes from bioinformaticians themselves, rather than from biologists. Moreover, this criticism does not seem to have had much effect on how bioinformatics is practiced, given the length of time over which it has been made.

For example, Carole Goble (2007. The seven deadly sins of bioinformatics. Keynote talk at the Bioinformatics Open Source Conference Special Interest Group at the 15th Annual International Conference on Intelligent Systems for Molecular Biology (ISMB 2007) in Vienna, July 2007) produced this list of what she called "intractable problems in bioinformatics":
1. Parochialism and insularity.
2. Exceptionalism.
3. Autonomy or death!
4. Vanity: pride and narcissism.
5. Monolith megalomania.
6. Scientific method sloth.
7. Instant gratification.
More recently, Manuel Corpas, Segun Fatumo & Reinhard Schneider (2012. How not to be a bioinformatician. Source Code for Biology and Medicine 7: 3) pointed out what they call "a series of disastrous practices in the bioinformatics field", which look very similar:
1. Stay low level at every level.
2. Be open source without being open.
3. Make tools that make no sense to biologists.
4. Do not provide a graphical user interface: command line is always more effective.
5. Make sure the output of your application is unreadable, unparseable and does not comply to any known standards.
6. Be unreachable and isolated.
7. Never maintain your databases, web services or any information that you may provide at any time.
8. Blindly believe in the predictions given, P-values or statistics.
9. Do not ever share your results and do not reuse.
10. Make your algorithm or analysis method irreproducible.
You can peruse the originals to check out the details of these problems, and whether they sound uncomfortably familiar.