3. In order to analyse phylogenetic relationship among a group of 20 organisms, you have selected a few
different sets of similar protein sequences. Every answer below characterizes a single set; provided are
ranges of lengths and percentage identities within the set. Select the set(s) suitable for phylogenetic
analysis.
a) 20 residues, 90-95% of identity
b) 200-300 residues, 50-70% of identity
c) 900-1000 residues, 70-90% of identity
d) 90-110 residues, 95-100% of identity
e) MSA of length 200 residues, edited to cut off
different columns
f) 200-300 residues, 10-20% of identity
14. The following statements were successfully executed usin Biopython:
from Bio.SeqIO import read
dna = read('some dna.fasta', 'fasta')
p1 = dna.seq.translate (table=1)
p2 = dna.seq.translate(table=2)
Knowing that the table argument of the Seq. translate method is used to specify the variant of the
genetic code, select the true sentence(s):
a) The dna represents a protein sequence.
b) p1 and p2 are handles, which can be used to
retrieve the corresponding protein sequences
from NCBI database.
c) p1 and p2 must be the same sequence,
because the protein code is not degenerated.
d) The sequences represented by p1 and p2 may
be the same.
e) From the context it is apparent, thet p1 and
p2 represent protein sequences.
f) p1 and p2 must be the same sequence,
because the protein alphabet is the same for
all organisms.
15. A student claims, that using DALI he found similarity between protein structures X and Y, with Ca
RMSD-3.5Ã… and sequence identity of 13%. What do you think of this result?
a) The RMSD value is way too low to suggest
d) The RMSD difference of 3.5 standard
structure similarity.
b) He is obviously lying, it is not possible to
compare structures with DALI.
deviations from the most probable result
suggest very low chance of similarity.
e) There is almost no chance of homology.
c) The sequences are misaligned and the value of f) The proteins are probably homologous.
RMSD is wrong.
16. Three students had to compare two protein structures. One protein had 100, the other 120 residues,
and their alignment had a 20-residue long gap in the middle of the shorter sequence.
Student A superimposed the first 100 Ca atoms of each structure and calculated the Ca RMSD for
them. Student B calculated RMSD between pairs of first 100 Ca atoms from each structure withou
superposition. Student C used the alignment to identify which Ca atoms are equivalent, and calculate
the superposition and then RMSD for these pairs. They obtained different results: 1) 19.5Ã…; 2) 6.7.
3) 2.3Ã….
Identify who calculated which number, and which one is correct.
a) A-1), B-2), C-3); correct is result 2)
d) A-3), B-1), C-2); correct is result 1)
b) A-3), B-1), C-2); correct is result 2)
e) A-3), B-2), C-1); correct is result 3)
c) A-3), B-2), C-1); correct is result 1)
f) A-2), B-1), C-3); correct is result 3)
17. You have isolated some unknown mRNA. It was reverse transcribed to cDNA, amplified with PC
sequenced. You have searched the Nucleotide database at NCBI for similar sequences with BL
but without significant results. Which service would be best to search for homologous gen
higher sensitivity?
a) PSI-BLAST
b) TBLASTX
c) SwissProt
d) MEGABLAST
e) BLASTP
f) BLASTX