Showing posts with label software. Show all posts
Showing posts with label software. Show all posts

Monday, March 9, 2020

A sneak peek into the upcoming SplitsTree 5


For some time now, the official SplitsTree page (www.splitstree.org) has been offline. The reason is that a major update is on the way: SplitsTree5. A beta version is already available, so let's take a quick look at it.

During installation you will be asked how much RAM you want to dedicate. Give as much as possible, in case you want to handle large tree with myriads of splits. I chose 16 GB (ie. half of the RAM installed on my PC).

Here's how it looks when you start the program:


The menus known from SplitsTree4 are still there, and the important functions appear to be already implemented. Some are new, and some have been moved:
  • Menu File: there is a new option is to "Export workflow", which produces a graphical representation (ie. a flow-chart) of what you did with the imported data, which is shown in the main display panel ("Workflow")
  • New menu Select: collects together the Select options formerly included under Edit.
  • The Trees option is now called Tree
  • Menu Network has all of the classics (distance-based phylogenetic networks, tree-based networks, character-based networks); but missing (so far) are the Pruned Quasi Median network and Spectral Splits options, possibly due to very little demand. An important new function is (or will be) that one can change between "Splits Network view" (ie. the view we are used to from SplitsTree4) and "Haplotype Network view" (as known from the TCS, NETWORK, etc. programs)
  • New menu PCoA, to do principal component analysis (at some point).
  • The menu Analysis appears to be still in development. Currently there are five options: Show Bootstrap tree..., Show Bootstrap network..., Estimate invariable sites..., Compute Phylogenetic Diversity, Compute Delta Score, and (new) Show workflow.
  • The menu Window will be split into Window and Help. Menu Help includes also now a direct link to the (new to me, and, noting the low number of discussion threads, apparently most of the world), a SplitsTree Community page (online since September 2017).
The new GUI reminds me a bit of RStudio —instead of pop-up windows vanishing once you perform a function, you will keep subsequent sheets in the panels. This makes it easier for new users.

When opening a data matrix not directly interpretable, you may activate the "Import" menu, asking you to specify the data type and the file format:


Eventually, as in SplitsTree4, the importer is currently sensitive to additional code and commentary brackets, and cannot eg. handle polymorphisms for categorical data (such as "(01)", "{01}"). Accordingly, importer warnings will pop up. Probably, a lot of testing and tweaking is required to make this work as planned. The selection list for file formats is comprehensive, but also ambitious. It may be a good idea to focus on a simple import format (eg. Phylip without its name-length restrictions, or clean NEXUS), and leave the import / export issues to other software packages (such as Mesquite, or R-conversion tools).

But we can read in Splits-NEXUS files generated by SplitsTree4 without any problems. To sneak a bit more:


A very nice function is that the flags in the analysis pipeline are fully interactive, allowing for quick manipulation / overview of what was used. For example, by clicking on "NeighborNet", we get a new panel for tweaking the NNets options or change the used method:


When moving above a menu item, a short explanation may pop up. The menus in the modification panel include drop-down boxes and input fields (here, for NNet):

Close-up of the NeighborNet panel.

Another important upgrade is the "Workflow" sheet, which gives you access to data filtering, methods and visualization etc., by just double-clicking on the respective item in the flow-chart (items can be dragged and moved, too):


Graphically, SplitsTree5 is functional as well. View > Format... (Ctrl-Shift-J) will open the remodeled coloring and type window in the method / lower left panel, where you can chose: font, label and (selected) edge(s) colors, node colors and shapes. In addition to circles and squares, we now have the choice between up- and down-triangles, diamonds, and hexagons. The graphical export option is gone (Ctrl-Shift-M; for now) and replaced by a modifiable, objects-containing PDF (similar to the ones produced by Dendroscope), generated simply by printing out to PDF.

The current beta version may not be able to fully replace SplitsTree4 yet (especially since the current manual only contains an 'Acknowledgments' section) but has already enough functionality (some new) to play around and explore the wonderful world of phylogenetic networks.

So, try it out for yourself.



Current issues

Glitches (on my Windows-PC running the latest Java version) that I have encountered include:
  • flickering scroll bars – but, when resizing the window a bit and keeping the left mouse button pressed, the flickering stops
  • I couldn't exit the program after opening more than one window / data set
  • a few menu items may not work yet (e.g. Select > All Labeled Nodes, Ctrl-Shift-L).
Moving edges can also be a bit tricky. You need to first select the edges, when the selected edge bundle will be highlighted by a broad yellow aura, and then move the pointer to one of the nodes, until the node is surrounded by an even broader aura. Then click and keep the mouse button down.

To get rid of node shapes, I had to click several times on "none" (first it changes to circles, which then become smaller until being nearly invisible).

Important note: While I had no problem in opening any of my SplitsTree4-generated and saved files, when saving a file in SplitsTree5, SplitsTree4 gives an import failure error message.

Tuesday, November 7, 2017

PhyloNetworks: a package for phylogenetic networks

Recently, another computer package was released that is of relevance to this blog. This is described in a forthcoming paper:
Claudia Solís-Lemus, Paul Bastide, Cécile Ané (2017) PhyloNetworks: a package for phylogenetic networks. Molecular Biology and Evolution (in press) 12: 3292-3298.
The authors describe the package this way:
PhyloNetworks is a Julia package for the inference, manipulation, visualization and use of phylogenetic networks in an interactive environment. Inference of phylogenetic networks is done with maximum pseudolikelihood from gene trees or multi-locus sequences (SNaQ), with possible bootstrap analysis. PhyloNetworks is the first software providing tools to summarize a set of networks (from a bootstrap or posterior sample) with measures of tree edge support, hybrid edge support, and hybrid node support. Networks can be used for phylogenetic comparative analysis of continuous traits, to estimate ancestral states or do a phylogenetic regression.


The  SNaQ analysis is described in a previous paper:
Solís-Lemus C, Ané C (2016) Inferring phylogenetic networks with maximum pseudolikelihood under incomplete lineage sorting. PLOS Genetics 12:e 1005896.
The phylogenetic model used incorporates: mutations (as usual), incomplete lineage sorting of alleles in ancestral populations (using the coalescent), and horizontal inheritance of genes (ie. reticulations in the network). The likelihood is decomposed into quartets, which makes the likelihood calculations relatively fast, and also allows the analyses to be scaled up to many species and many genes.

The PhyloNetworks software is open source, and is available with documentation at:
https://github.com/crsl4/PhyloNetworks.jl
Have fun learning to use the Julia system, which I had never even heard of before investigating this new package!

Note: In spite of the similarity in name, this new package has nothing to do with Luay Nakhleh's PhyloNet package, nor to the Phylogenetic Networks blog.

Tuesday, September 5, 2017

SPECTRE: a suite of phylogenetic tools for reticulate evolution


Recently, the Earlham Institute, in the UK, released a set of software tools that are of relevance to this blog — SPECTRE. These tools are described in a forthcoming paper:
Sarah Bastkowski, Daniel Mapleson, Andreas Spillner, Taoyang Wu, Monika Balvočiūte and Vincent Moulton (2017) SPECTRE: a Suite of PhylogEnetiC Tools for Reticulate Evolution. [Now published.]

This is a toolkit rather than simple-to-use program, meaning that the various analyses exist as separate entities that can be combined in any way you like. More importantly, new analyses can be added easily, by those who want to write them, which is not the case for more commonly used programs like SplitsTree. This way, the analyses can also be incorporated into processing pipelines, rather than only being used interactively.

Apart from the usual access to data files (including Nexus, Phylip, Newick, Emboss and FastA formats), the following network analyses are currently available:
NeighborNet, NetMake, QNet, SuperQ, FlatNJ, NetME
The program also outputs the networks, of course. Here is an example of the SPECTRE equivalent of a NeighborNet analysis from a recent blog post (where the network was produced by SplitsTree, and then colored by me).


Running the program(s) is relatively straightforward, once you get things installed. Installation packages are available for OSX, Windows and Linux.

Sadly, for me installation was tricky, because SPECTRE requires Java v.8, which is unfortunately not available for OSX 10.6 (which runs on most of my computers). Even getting Java v.8 installed on the one computer I have with a later version of OSX was not easy, because installing a Java Runtime Environment (the JRE download file) from Oracle does not update the Java -version symlinks or add Java to the software path — for this I had to install the full Java Development Kit (the JDK download file). Sometimes, I hate computers!

Monday, December 9, 2013

Results of some bioinformatics polls


In 2008, Michael Barton conducted a Bioinformatics Career Survey. Since then, various groups have updated some of that information by conducting polls of their own. Below, I have included some of the more recent results, for your edification.

This first one comes from the Bioinformatics Organization, in response to the question: What is your undergraduate degree in? It is interesting to note that more bioinformaticians are biologists by training, rather than computational people.


The next one is actually an ongoing poll at BioCode's Notes, in response to the question: Which are the best programming languages for a bioinformatician? R is an interesting choice as the most useful language, given the more "traditional" use of Perl and Python.


That leads logically to another of the Bioinformatics Organization's questions: Which computer language are you most interested in learning (next) for bioinformatics R&D? I guess that if you already know R, then either Python or Perl is a useful thing to learn next.


Furthermore, the Bioinformatics Organization also asked: Which math / statistics language / application do you most frequently use? The choice of R here is more obvious, given that it is free, which most of the others are not. I wonder what the answer "none of the above" refers to.


Wednesday, November 20, 2013

Bioinformaticians look at bioinformatics


Bioinformatics as a term dates back to the 1970s, usually credited to Paulien Hogeweg, of the Bioinformatics group at Utrecht University, in The Netherlands, although it apparently did not make it into print until 1988 (Paulien Hogeweg. 1988. MIRROR beyond MIRROR, puddles of Life. In: Artificial Life, C. Langton, ed. Addison Wesley, pp. 297-315.).

In the 1990s the field expanded rapidly and became recognized as a discipline of its own, as a subset of computational science. However, Christos A. Ouzounis (2012. Rise and demise of bioinformatics? Promise and progress. PLoS Computational Biology 8: e1002487) has noted a distinct decrease in the use of the term itself, as shown by this graph.


Ouzounis recognizes three (admittedly artificial) periods in the history: Infancy (1996-2001), Adolescence (2002-2006) and Adulthood (2007-2011). Along the way, the practice of bioinformatics has received a lot of criticism. I have noted some of this before, in previous blog posts:
Poor bioinformatics?
Archiving of bioinformatics software

What is perhaps most important is that much of this criticism comes from bioinformaticians themselves, rather than from biologists. Moreover, this criticism does not seem to have had much effect on how bioinformatics is practiced, given the length of time over which it has been made.

For example, Carole Goble (2007. The seven deadly sins of bioinformatics. Keynote talk at the Bioinformatics Open Source Conference Special Interest Group at the 15th Annual International Conference on Intelligent Systems for Molecular Biology (ISMB 2007) in Vienna, July 2007) produced this list of what she called "intractable problems in bioinformatics":
1. Parochialism and insularity.
2. Exceptionalism.
3. Autonomy or death!
4. Vanity: pride and narcissism.
5. Monolith megalomania.
6. Scientific method sloth.
7. Instant gratification.
More recently, Manuel Corpas, Segun Fatumo & Reinhard Schneider (2012. How not to be a bioinformatician. Source Code for Biology and Medicine 7: 3) pointed out what they call "a series of disastrous practices in the bioinformatics field", which look very similar:
1. Stay low level at every level.
2. Be open source without being open.
3. Make tools that make no sense to biologists.
4. Do not provide a graphical user interface: command line is always more effective.
5. Make sure the output of your application is unreadable, unparseable and does not comply to any known standards.
6. Be unreachable and isolated.
7. Never maintain your databases, web services or any information that you may provide at any time.
8. Blindly believe in the predictions given, P-values or statistics.
9. Do not ever share your results and do not reuse.
10. Make your algorithm or analysis method irreproducible.
You can peruse the originals to check out the details of these problems, and whether they sound uncomfortably familiar.

Wednesday, July 3, 2013

Archiving of bioinformatics software


Some months ago I wrote a blog post about what is perceived to be the rather poor quality of many computer programs in bioinformatics (Poor bioinformatics?), noting that many bioinformaticians aren't taking seriously the need to properly engineer software, with full documentation and standard programming development and versioning.

An obvious follow-up to that post is to consider the archiving of bioinformatics software. If programs are written well, then they should be permanently archived for future reference. A number of bloggers have commented on what is perceived to be the poor current state of affairs here, as well, and I thought that I might draw your attention to a few of the posts.

In many ways, this issue is the computational equivalent of storing biological data, about which I have also written recently (Releasing phylogenetic data). My comments about this were:
  • There is a difference between storing / releasing the original data (eg. raw DNA sequences) and the data as analyzed (eg. aligned sequences)
  • There are sustainable and accessible archiving facilities for raw data that are almost universally used (eg. GenBank)
  • Many people do not release the processed data as analyzed (some of them will if directly asked to do so)
  • Many of the people who do release their analyzed data do so on the homepage of one of the authors, which is better than nothing but is rarely sustainable
  • There are sustainable and accessible archiving facilities for processed data, such as TreeBASE and Dryad.
Analogous comments can be made about the archiving of bioinformatics software.

The first question to ask is this: what proportion of the bioinformatics software referred to in publications is actually stored in sustainable and accessible archives? A corollary to this question is: what archive facilities are being used? Casey Bergman, at the I Wish You'd Made Me Angry Earlier blog, has attempted to answer both of these questions (Where Do Bioinformaticians Host Their Code?).

In answer to the first question, Casey notes:
of the many thousands of articles published in the field of bioinformatics, as of Dec 31 2012 just under 700 papers (n=676) have easily discoverable code linked to a major repository in their abstract.
While many papers may have the code URL in the Methods or Results sections but not the Abstract, this does suggest that repository archiving is not the mode actually employed by bioinformaticians. Instead, they are archiving (if at all) on personal or institutional homepages.

Sadly, the reported rate of decay of URLs ("Error 404: Page not found") indicates that this is rarely a sustainable approach to archiving (eg. see the Google+ comment by Dave Lunt). The relevance of the similar situation with the TreeBase / Dryad type of repository has not gone unnoticed, for example by Hilmar Lapp. These repositories require and enforce standards of data and software archiving, as well as providing persistence.

The answer to the second question, about which repositories, seems to be (see also the data provided by MRR in the comments to Casey's blog post):
  • SourceForge has been vastly predominant
  • Google Code has a large number of projects, but many of them have never made it to publication
  • GitHub has had a rapid recent growth rate, and therefore appears to be becoming the preferred repository.
Other repositories, such as BitBucket, seem to be much less used. Users on other forums (eg. Biostar: Where would you host your open source code repository today?) seem to concur with the choice of GitHub, mainly because of the tools available (it is user-oriented rather than project-oriented).

This leads to the issue of how permanent the archiving is at the major repositories. It turns out that there is a major difference in policies, as noted by Casey Bergman:
SourceForge has a very draconian policy when it come to deleting projects, which prevents accidental or willful deletion of a repository. In my opinion, Google Code and (especially) GitHub are too permissive in terms of allowing projects to be deleted.
In a follow-up post (On the Preservation of Published Bioinformatics Code on GitHub), Casey expands on this theme:
A clear trend emerging in the bioinformatics community is to use GitHub as the primary repository of bioinformatics code in published papers. While I am a big fan of Github and I support its widespread adoption, I have concerns about the ease with which an individual can delete a published repository. In contrast to SourceForge, where it is extremely difficult to delete a repository once files have been released, and this can only be done by SourceForge itself, deleting a repository on GitHub takes only a few seconds and can be done (accidentally or intentionally) by the user who created the repository.
This is an important issue, as exemplified by Christopher Hogue in the comments section of that blog post:
In my case SourceForge preserved the SLRI toolkit my group made in Toronto. As the intellectual property underlying the code was sold to Thompson-Reuters in 2007, my host institution and the dealmakers pressured me to delete the repository. SourceForge policy kept it on the site ... [However,] the aftermath of all this is that, of everything my group did under the guise of open source, only about 30% is preserved and online, and the rest is buried in an intellectual property shoebox at Thompson-Reuters. Host institutions have a lot of power of ownership over your intellectual property. If you win the right to post work into open-source, the GitHub delete policy means that your host institution can over-ride this, and require you to take your code out of circulation. GitHub is great, but for the sake of preservation, SourceForge has the right policy, protecting your decision to go open source from later manipulations by your host institution when it becomes "valuable".
Casey Bergman's response to this issue has been to create the Bioinformatics Archive on GitHub. This is based on the idea used by the journal Computers & Geosciences, in which the journal editor forks the GitHub code into a journal "organization" for all accepted papers — this creates a permanent repository, which is necessary because deleting a private GitHub repository will delete all forks of the repository but deleting a public repository will not do so. So, Casey has been personally forking the code for all publications that come to hand (currently 147 repositories) into the Bioinformatics Archive, thus creating a public repository for all of the relevant GitHub code.

However, this is clearly a stopgap measure. Dave Lunt, at the EvoPhylo blog, has listed three desiderata for a more permanent solution to the issue (How can we ensure the persistence of analysis software?):
  • A publisher driven version of the Bioinformatics Archive; journals should have a policy for the hosting of published code in a sustainable and accessible archive in a standardized manner
  • Redundancy to ensure persistence in the worst case scenario; archive persistence is the key requirement, and this can only happen in public repositories, with the published URL and/or DOI pointing to a public copy of the code
  • The community to initiate actual action; authors need to pressure the publishers to adopt a Dryad-like strategy, in which a large group of ecology and evolutionary biology journals agreed to require the use of a public database for storing the biological data associated with their publications.

At a minimum, a persistent public repository is a snapshot of the code at the time of publication, just as a sequence alignment is a snapshot of the processed data at the time of its publication. This does not preclude further work on the code, and further publications based on the newly modified code, just as new sequence alignments can be created by adding newly acquired sequences. Open-source code can still be newly forked, and there can be user-contributed updates and public issue tracking. Multiple snapshots of code related to different publications through time is not necessarily an issue, but it will need to be handled in some sensible manner.

The main reason for requiring the public archiving of code is to deal with the all-too-common situation when code is no longer being maintained (the scholarship ran out, the grant ended, the author retired, etc). For example, Jamie Cuticchia & Gregg Silk (2004, Bioinformatics needs a software archive, Nature 429: 241) mention the loss of part of the code when the multi-million dollar Genome Database lost funding in 1998. These two authors seem to be the first to have proposed a Bioinformatics Software Archive, "in which an archival copy of bioinformatics software would be maintained in a secure central repository supported by public funding." Personal and institutional homepages are too ephemeral (suffering what is known as URL decay) and too prone to politics to be considered acceptable for the storage of data and software in high-quality science.

Saturday, September 8, 2012

Poor bioinformatics?


There have been a number of recent posts in the blogsphere about what is perceived to be the rather poor quality of many computer programs in bioinformatics. Basically, many bioinformaticians aren't taking seriously the need to properly engineer software, with full documentation and standard programming development and versioning.

I thought that I might draw your attention to a few of the posts here, for those of you who write code. Most of the posts have a long series of comments, which are themselves worth reading, along with the original post.

At the Byte Size Biology blog, Iddo Friedberg discusses the nature of disposable programs in research, which are written for one specific purpose and then effectively thrown away:
Can we make accountable research software?
Such programs are "not done with the purpose of being robust, or reusable, or long-lived in development and versioning repositories." I have much sympathy for this point of view, since all of my own programs are of this throw-away sort.

However, Deepak Singh, at the Business|Bytes|Genes|Molecules blog, fails to see much point to this sort of programming:
Research code
He argues that disposable code creates a "technical debt" from which the programmer will not recover.

Titus Brown, at the Living in an Ivory Basement blog, extends the discussion by considering the consequences of publishing (or not publishing) this sort of code:
Anecdotal science
He considers that failure to properly document and release computer code makes the work anecdotal bioinformatics rather than computational science. He laments the pressure to publish code that is not yet ready for prime-time, and the fact that computational work is treated as secondary to the experimental work. Having myself encountered this latter attitude from experimental biologists (the experiment gets two pages of description and the data analysis gets two lines), I entirely agree. Titus concludes with this telling comment: "I would never recommend a bioinformatics analysis position to anyone — it leads to computational science driven by biologists, which is often something we call 'bad science'." Indeed, indeed.

Back at the Business|Bytes|Genes|Molecules blog, Deepak Singh also agrees that "a lot of computational science, at least in the life sciences, is very anecdoctal and suffers from a lack of computational rigor, and there is an opaqueness that makes science difficult to reproduce":
Titus has a point

This leads to Iddo Friedberg's post (the Byte Size Biology blog) that mentions this concept:
The Bioinformatics Testing Consortium
This group intends to act as testers for bioinformatics software, providing a means to validate the quality of the code. This is a good, if somewhat ambitious, idea.


Finally, Dave Lunt, at the EvoPhylo blog, takes this to the next step, by considering the direct effect on the reproducibility of scientific research:
Reproducible Research in Phylogenetics
He notes that bioinformatics workflows are often complex, using pipelines to tie together a wide range of programs. This makes the data analysis difficult to reproduce if it needs to be done manually. Hence, he champions "pipelines, workflows and/or script-based automation", with the code made available as part of the Methods section of publications.

Tuesday, June 12, 2012

New version of phylogenetic networks software

There's a new release of Dendroscope with several new methods for constructing phylogenetic networks. This blog will explain the new functionality added to the Cass algorithm.

Cass can be used to construct a rooted phylogenetic network from any number of rooted trees. These trees can be multifurcating (nonbinary). The output network will display all clusters of the input trees. In certain situations the algorithm has been shown to minimize the reticulation number of the network (see van Iersel et al. 2010 and Kelk et al. 2012).

The new release of Cass can also produce networks that display the input trees (rather than only the clusters from the input trees). In other words, Cass can be used as a heuristic for the problem Hybridization Number on any number of multifurcating trees. This is a notoriously hard problem and at the moment Cass is the only implemented algorithm for this problem.

For inputs consisting of two multifurcating trees, Cass solves Hybridization Number optimally (Kelk et al. 2012) and, for such instances, Dendroscope also contains a faster optimal algorithm (go to Algorithms - Hybridization Networks) by Huson and Linz (submitted). For bifurcating trees, there is one other heuristic available that can be used for any number of trees: the program PIRN by Yufeng Wu.

But, as I said, for more than two multifurcating trees, Cass is the only method currently available.

Let's have a look at the new functionality of Cass for these input trees:



The Cass algorithm can be found here (Algorithms - Level-k Network Consensus):


You get several new options:


If we don't check the box "construct only networks that display the trees", we get the following network with one reticulation, which displays all clusters of all input trees:


If we do check the box "construct only networks that display the trees", we get the following three networks with two reticulations each. Each of the networks displays all three input trees.


You see that in this case one needs more reticulations to display the trees than to display the clusters from the trees. You also see that Cass can now produce several solutions rather than just one. Note however that Cass is not guaranteed to find all optimal solutions. As a result of a collapsing step in Cass, it might miss networks (see Kelk et al. 2012). That is also the reason why Cass is not guaranteed to find an optimal solution (unless the input consists of only two trees or the output network is at most level-2).

Finally, Cass can also be used to construct a network from a multi-labelled tree. Go to Algorithms - Multi-Labelled-Tree to Network - MUL to Network, level-k-based.