Genbank accession
URC25646.1 [GenBank]
Protein name
central tail fiber J
RBP type
TF
Evidence Phold
Probability 1,00
TSP
Evidence RBPdetect
Probability 0,83
TF
Evidence RBPdetect2
Probability 0,96
Protein sequence
MAKYMISGSKGGSKKPYVPKEMEDNLISINKIKVLLAVSDGECDPDFTLRDLYLDDVPVIASDGTVNYEGVTAEYRPGTQTQDYIQGFTDTSSEVTVARDITGDNPYVISVTNKNLSAVRIKILMPVGIKTEDNGDLVGVRVEYAVDMAVDGGSYSEVMRDVIDGKTRSGYDRSRRIDLPKFDERVLIRVKRLTPDSTSSKVTDKIKLQSYAEVVDAKFRYPLTGLVFVEFDSELFPTQIPNISIKKKWKIINVPSNYDPISREYHGSWDGTFKKAWSNNPAWVLYDLVTNQRYGLDQRELGIQIDKWSLYEAGVYCDQKVPDGKGGTEPRYLCDVVIQNQVEAYQLIRDICSIFRGMSFWNGESLSIVIDKPRDPSYVFTNENVINGDFQYTNASEKSMYTQCNVTFDDEQNMYQQDVEGVFDTEAALRFGYNPTSITAIGCTRRSEANRRGRWVLKTNLRSTTVNFATGLEGMIPSIGDVIAIADNFQSSNLTLNLSGRVMEVSGLQVFVPFKVDARPGDFIIINKPDGKPVKRTISKVSADGKTIELNIGFGFDVNPDTVFAIDRTDLALQQYVVTTISKGDDENEFTYSITAVEYDPNKYDEIDYGVNIDDRPTSIVQPDVMAAPENVKISSYSRVVQGVSVETMVVSWDKVPYASLYEMQWRKGDGNWLNTPQTANKEIEVEGIYSGNYQVRVRSVSASGNASPWSKIATATLTGKVGEPGAPINLTASDNEVFGIRVKWGMPEGSGDTAYIELHQSPDGTVENSSLLTLIPYPQYEYWHSTLPAGQVVWYRIRSVDRIGNVSSWTDFVRGMASDDVESVLGDILDKIFDTEAGQEIKENAIDSANKIKDQAQSIIQNALANDADVKWTRVQNGKRKAEYGHALELIANETEARVTQIEELRASIDGEITSSIKTVQEAIATESETRATQIQQLDSKFTKEIDGVRKDTSASISDVRQTITNESEARAQAVQQLDAKFTKEINDLDGVIKTEVEANISEVKQAIANETEARVQADQALTARFGDVESALVEKLDSWASVDSVGAKYAMKLGLTYKGQQYSAGMVMQLSQGSSGLISQILFDANRFAIMTSSTGGTFTLPFVVENNQVFINSLLVKNGSITNAMIGNVIQSNNFVQNQQGWRLDKNGIFENYGSTPGEGATKFTNEGLKVKDANGVLRVEVGRITGSW
Physico‐chemical
properties
protein length:1192 AA
molecular weight: 132089,32580 Da
isoelectric point:4,80371
aromaticity:0,08557
hydropathy:-0,37366

Domains

Domains [InterPro]
DC_0014
STR
1–965
IPR053171
Unmapped
2–872
IPR013783
STR
627–721
IPR003961
STR
628–717
IPR003961
STR
628–722
URC25646.1
1 1192
Architecture
STR
ATT
STR
ATT
STR
RBD
STR 1-91 | ATT 92-217 | STR 218-339 | ATT 340-491 | STR 492-965 | RBD 966-1191 |
Legend: ATT STR RBD CBM LEC ENZ CHP LNK TAS TTP UNK Unmapped

Tail Spike Domain Segmentation

Tail Spike Domain Segmentation

This protein has been segmented into three structural domains: N-terminal, central domain, and C-terminal.

Domain Layout
N-terminal
Central
C-terminal
URC25646.1
1 1192
Domain Start End Length (AA) Confidence
N-terminal 1 507 507 0,9280
Central domain 508 1067 561 0,0439
C-terminal 1068 1192 124 0,3902
Legend: N-terminal Central domain C-terminal
3D Structure with Domain Coloring

The structure is colored according to the domain segmentation: N-terminal (blue), Central (green), C-terminal (pink).

Domain Coloring
N-terminal
1-507
Central
508-1067
C-terminal
1068-1192

Taxonomy

  Name Taxonomy ID Lineage
Phage Escherichia phage EC195
[NCBI]
2936946 Viruses > Duplodnaviria > Heunggongvirae > Uroviricota > Caudoviricetes
Host Escherichia coli
[NCBI]
562 cellular organisms > Bacteria > Pseudomonadati > Pseudomonadota > Gammaproteobacteria > Enterobacterales

Coding sequence (CDS)

Coding sequence (CDS)
Genbank protein accession
URC25646.1 [NCBI]
Genbank nucleotide accession
ON185588.1 [NCBI]
CDS location
range 16414 -> 19992
strand +
CDS
ATGGCTAAATATATGATAAGCGGCAGTAAGGGCGGAAGCAAAAAGCCATACGTGCCAAAAGAGATGGAAGATAACCTGATCTCGATAAACAAGATTAAAGTTTTGCTGGCTGTATCTGATGGCGAGTGCGATCCAGATTTCACGTTGCGCGATCTTTATCTTGATGATGTTCCGGTTATTGCCAGCGATGGCACTGTTAACTACGAGGGTGTTACTGCTGAATATCGACCAGGTACGCAGACGCAAGATTACATCCAGGGGTTTACTGACACATCAAGCGAGGTGACAGTTGCAAGAGATATTACCGGAGACAATCCTTATGTTATTTCTGTTACAAATAAAAATCTATCTGCGGTAAGAATAAAGATCCTGATGCCAGTAGGCATTAAAACAGAGGATAACGGCGATCTTGTTGGCGTAAGGGTTGAGTATGCCGTAGATATGGCTGTTGATGGCGGTTCTTATAGCGAGGTTATGAGAGATGTAATTGACGGCAAGACAAGATCAGGATACGACCGCAGCAGAAGGATTGATCTTCCTAAGTTTGATGAGCGCGTTTTAATCCGAGTAAAGCGACTGACTCCAGACAGCACATCTTCAAAGGTGACTGATAAAATCAAGCTGCAAAGTTACGCTGAGGTTGTTGATGCAAAATTCCGCTATCCTCTGACTGGACTTGTATTCGTAGAATTTGACAGCGAATTGTTTCCTACGCAAATCCCTAACATTTCTATAAAAAAGAAATGGAAGATTATTAATGTGCCAAGCAACTATGATCCAATATCAAGAGAATATCACGGGTCATGGGATGGGACTTTTAAAAAAGCGTGGTCAAATAATCCTGCTTGGGTTCTTTATGATCTGGTGACAAATCAGCGTTATGGACTTGATCAGCGAGAGTTAGGAATACAGATCGACAAGTGGAGCTTATACGAGGCTGGCGTTTACTGCGATCAGAAAGTTCCAGACGGTAAAGGCGGCACTGAGCCTCGCTACCTATGCGATGTGGTGATTCAGAATCAAGTTGAGGCTTATCAGCTAATCCGTGACATTTGCTCAATCTTTCGCGGAATGAGTTTTTGGAATGGTGAGAGCTTATCAATCGTGATTGATAAGCCGCGCGATCCATCATACGTGTTTACTAATGAAAACGTCATCAACGGTGATTTTCAGTACACAAACGCAAGCGAAAAAAGCATGTACACGCAGTGTAACGTGACGTTTGACGACGAACAAAACATGTATCAGCAGGACGTAGAGGGGGTTTTTGATACTGAGGCGGCATTACGATTTGGATACAACCCAACAAGCATTACAGCGATCGGGTGTACGCGCAGGAGCGAAGCGAATCGTCGCGGTCGTTGGGTTTTGAAAACAAACCTTAGAAGCACTACTGTAAACTTTGCTACCGGACTAGAGGGGATGATTCCATCAATAGGTGATGTTATTGCTATCGCTGATAATTTTCAGAGTAGCAACCTAACGTTAAACCTATCGGGCCGAGTAATGGAAGTTTCAGGGCTGCAGGTTTTCGTTCCGTTTAAGGTTGATGCTCGCCCTGGTGATTTTATTATCATCAACAAGCCGGACGGCAAGCCAGTTAAGCGCACGATCTCAAAGGTTAGCGCAGACGGAAAAACCATTGAGTTAAATATTGGATTTGGTTTTGATGTTAATCCTGATACTGTTTTTGCGATTGACCGTACTGATCTTGCGTTGCAGCAATACGTTGTGACAACCATCAGCAAGGGTGATGACGAAAACGAGTTTACCTATTCAATCACGGCTGTAGAGTACGATCCGAACAAATACGACGAGATTGATTATGGAGTAAACATTGATGACAGACCGACTTCAATTGTTCAGCCTGACGTGATGGCAGCGCCTGAGAACGTTAAGATCTCATCTTATTCTCGCGTCGTGCAGGGTGTTAGCGTTGAGACTATGGTTGTTTCATGGGATAAGGTTCCTTACGCATCGCTTTATGAAATGCAGTGGCGAAAAGGTGATGGTAACTGGCTGAATACGCCGCAGACCGCTAACAAAGAGATAGAGGTAGAAGGGATTTACTCTGGCAACTACCAAGTAAGGGTGAGATCCGTTTCTGCAAGCGGTAACGCTTCCCCGTGGTCAAAGATTGCAACCGCCACTCTGACAGGTAAAGTTGGCGAGCCAGGAGCGCCGATTAATCTTACGGCTTCTGATAATGAAGTTTTTGGCATTCGTGTCAAATGGGGTATGCCGGAAGGATCAGGCGATACGGCTTACATTGAGCTTCACCAATCGCCAGACGGAACGGTTGAAAACTCAAGTCTGCTTACGCTGATTCCATATCCTCAATATGAGTATTGGCATAGCACGTTACCAGCGGGGCAAGTTGTATGGTATAGAATCCGCAGCGTTGACAGAATAGGCAACGTTTCAAGCTGGACTGACTTTGTTCGCGGCATGGCGTCAGATGACGTTGAATCTGTTTTGGGCGACATTCTGGACAAGATTTTTGATACAGAAGCTGGTCAAGAAATCAAAGAGAACGCCATAGACAGTGCCAATAAAATCAAAGACCAGGCGCAATCAATCATACAGAACGCGTTGGCAAATGATGCAGATGTGAAGTGGACGCGAGTGCAAAACGGAAAGCGCAAGGCTGAATATGGTCATGCTCTTGAGCTTATCGCCAATGAAACAGAAGCGCGCGTAACTCAAATCGAAGAGTTAAGGGCTTCAATTGATGGCGAGATAACATCAAGCATCAAGACAGTGCAGGAGGCAATTGCCACTGAATCAGAGACGCGAGCGACTCAAATTCAGCAGCTTGATTCTAAATTCACAAAAGAAATCGACGGCGTGCGCAAGGATACTTCTGCAAGCATTAGCGATGTAAGGCAGACAATCACTAACGAGTCAGAAGCGCGCGCTCAGGCCGTTCAGCAGCTTGACGCTAAGTTCACGAAAGAGATAAACGACCTTGACGGAGTTATCAAAACAGAAGTCGAGGCTAACATCTCAGAAGTGAAACAGGCGATCGCCAATGAGACAGAGGCAAGGGTTCAGGCTGACCAGGCATTAACAGCACGATTTGGCGACGTTGAATCTGCATTGGTTGAAAAGTTGGATTCTTGGGCGAGCGTTGATTCAGTTGGGGCTAAATACGCTATGAAACTTGGCCTTACTTACAAAGGCCAGCAATACAGCGCAGGAATGGTGATGCAGCTTTCGCAGGGTTCATCCGGCCTTATCTCGCAAATTTTGTTTGATGCTAACAGGTTCGCCATTATGACTAGCTCTACTGGAGGGACTTTTACTTTGCCTTTCGTGGTTGAGAATAATCAGGTTTTCATTAATAGTCTTTTGGTGAAGAACGGTTCAATCACTAATGCGATGATTGGTAATGTGATTCAGTCAAACAACTTTGTTCAAAACCAGCAAGGATGGAGGCTTGATAAAAACGGAATCTTTGAGAATTACGGATCAACGCCAGGAGAAGGAGCTACTAAATTCACCAATGAGGGATTGAAGGTAAAAGATGCAAACGGAGTATTGAGGGTTGAAGTCGGAAGGATTACCGGAAGCTGGTAA

Genome Context

Genome Context

Tertiary structure

PDB ID
17b594f92469d1ec574bbdda1132e767b21181b269fb03e5378354c1359b3c0c
ESMFold
Source ESMFold
Method ESMFold
Resolution 0,7726
Oligomeric State monomer
Model Confidence
Very high
pLDDT > 90
High
90 > pLDDT > 70
Low
70 > pLDDT > 50
Very low
pLDDT < 50

Literature

Title Authors Date PMID Source
Complete genome sequences of 17 Escherichia coli bacteriophages isolated from wastewater, pond water, cow manure and bird feces Vitt,A.R., Ahern,S.J., Gambino,M., Holst Sorensen,M.C. and Brondsted,L. 2022-10-20 GenBank