Showing posts with label NCBI. Show all posts
Showing posts with label NCBI. Show all posts

Thursday, August 30, 2012

NCBI. Why? Why? Why is the default database for blasting "human G+T"

It is the little things.  The little things that can sometimes eat at you.  And here is one of my little pet peeves.  At some point - not sure how recently - NCBI changed the default database for Blastn searching to the "human G+T" database.  This Db contains human genomic and transcriptomic data.

Now - I am extremely grateful for many - if not most - of the things NCBI does.  Pubmed is great.  Pubmed Central rocks.  Sequence databases galore.  And all sorts of widgets and tools associated with sequence analysis.  Including a free blast server.  Which I use a lot.  But why? why is human G+T the default database for blastn searches?  Why?  I never only want to search this Db.  Why not have the non redundant Db of everything be the default?  Are there really that many people who only want to search human G+T?  Or is this some ploy to force people to do such searches?  This tiny little thing.  This setting.  This glitch.  Whatever it is.  It drives me batty ...

Tuesday, February 08, 2011

Though I generally love NCBI, the Sequence/Short Read Archive (SRA) seems to have issues; what do others think?

Well, here goes. Hope to not get people from NCBI too pissed off here. Overall, I think NCBI is invaluable: GenBank. PubMed. PubMed Central (PMC) (well, I have some complaints about that but let's not get into those here -- I still like it), BLAST (Basic Local Alignment Search Tool) and a plethora of other tools, databases and resources. Generally, money well spent.

However, one database from NCBI is driving me a bit wacky these days. This is the Sequence Read Archive (SRA). Known to some as the "Short Read Archive" this database is supposedly for storing "sequencing data from the next generation of sequencing platforms including Roche 454 GS System®, Illumina Genome Analyzer®, Life Technologies AB SOLiD System® , Helicos Biosciences Heliscope®;, Complete Genomics®, and Pacific Biosciences SMRT®."

It certainly seems to be used for that function. But alas, storing sequence is not the only need here. Recovering sequence and making use of it is really the key. And this is the area I have been having trouble with (especially related to environmental studies like rRNA PCR and metagenomics). Rather than go on about my particular issues here (and thus possibly biasing the discussion too much), I am wondering what others think of the SRA? Usability? Ease of deposition? Ease of extraction? Missing features? Things it does or does not do well? Do we need a new system for environmental projects?

Any and all comments welcome here or on twitter or on Friendfeed or wherever. See Friendfeed stream below:




Here are some comments so far from twitter
  • digitalbio Sandra Porter I agree. RT @phylogenomics: Though I generally love NCBI, the Sequence/Short Read Archive (SRA) seems to hav… (cont) http://deck.ly/~XM75A
  • lswenson Luke Swenson @phylogenomics I was JUST trying to navigate the SRA! There's no help section to be found, and forget about depositing sequences!
  • audyyy Davis-Richardson @phylogenomics I can never tell if my submission went through without emailing support. Also, no FASTQ support?
  • cabbageRed Rich C .@phylogenomics I agree, the SRA doesn't seem to be the easiest repository to search with what I believe to be "typical" NGS queries

Most recent post

A ton to be thankful for -- here is one part of that - all the acknowledgement sections from my scholarly papers

So - it is another Thanksgiving Day and in addition to thinking about family, and football, and Alice's Restaurant, I also think a lot a...