|
Name |
Accession |
Description |
Interval |
E-value |
| NHL_like_1 |
cd14953 |
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ... |
874-1226 |
1.81e-40 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271323 [Multi-domain] Cd Length: 323 Bit Score: 153.45 E-value: 1.81e-40
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 874 VSSIMGNGRRrsiscpscnGQADGNKLLA----PVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVL----ELSSN--- 940
Cdd:cd14953 1 VSTVAGSGTA---------GFSGGGGTAArfnsPSGVAVDAAGNLYVADRgnHRIRKITPDGVVTTVAgtgtAGFADggg 71
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 941 PAHRYY----LATDPvTGDLYVSDTNTRRIYRpksLTGAKDLTknaeVVAGTGEqclpfdeARCGDGGKAVEATLMSPKG 1016
Cdd:cd14953 72 AAAQFNtpsgVAVDA-AGNLYVADTGNHRIRK---ITPDGVVS----TLAGTGT-------AGFSDDGGATAAQFNYPTG 136
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1017 MAVDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLTSAR--PLTcdtsmhisQVRLEWPTDLAINPMDNsIYVLD--N 1090
Cdd:cd14953 137 VAVDAAGNLYVADTGnhRIRKITPDGVVTTVAGTGGAGYAGdgPAT--------AAQFNNPTGVAVDAAGN-LYVADrgN 207
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1091 NVVLQITENRQVRIAAGRPmhcqvpGVEYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVA 1170
Cdd:cd14953 208 HRIRKITPDGVVTTVAGTG------TAGFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGN---HRIRKITPAGVVTTVA 278
|
330 340 350 360 370
....*....|....*....|....*....|....*....|....*....|....*..
gi 2421800221 1171 GIPSEcdckndancdcyQSGD-GYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAV 1226
Cdd:cd14953 279 GGGAG------------FSGDgGPATSAQFNNPTGVAVDAAGNLYVADTGNNRIRKI 323
|
|
| Tox-GHH |
pfam15636 |
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH ... |
2343-2420 |
6.94e-40 |
|
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH/Endonuclease VII fold present in bacterial polymorphic toxin systems with a characteriztic sG[HQ]H signature motif. In bacterial polymorphic toxin systems, the toxin is exported by the type 2, type 6, type 7 or TcdB/TcaC-type secretion system. The metazoan teneurin proteins possess an inactive of this domain at their C-terminus.
Pssm-ID: 464783 Cd Length: 78 Bit Score: 142.75 E-value: 6.94e-40
10 20 30 40 50 60 70
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 2421800221 2343 EEKARILEQARQRALARAWAREQQRVRDGEEGARLWTEGEKRQLLSAGKVQGYDGYYVLSVEQYPELADSANNIQFLR 2420
Cdd:pfam15636 1 EERKRLLEHAKKRAVREAWHRERQLLRNGLPGSRDWTDEEKEELLSTGSVPGYDGEYIHPVEQYPELADDPSNIRFRK 78
|
|
| RhsA |
COG3209 |
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction ... |
1180-2120 |
1.76e-33 |
|
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction only];
Pssm-ID: 442442 [Multi-domain] Cd Length: 1103 Bit Score: 142.20 E-value: 1.76e-33
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1180 NDANCDCYQSGDGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKNKPLLNSMNFYEVASPTDQELYIFDINGTHQ 1259
Cdd:COG3209 109 AAATASAGRLVSTGAGAGGTVTAATGGTLGATAGSATTGSTDGGRGGVAVTGLAGGGASAYGLTLGGAAAGPATGVGTGA 188
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1260 YTVSLVTGDYLYNFSYSNDNDITAVTDSNGNTLRIRRDPNRMPVRVVSPDNQVIWLTIGTNGCLKSMTAQGLELVLFTYH 1339
Cdd:COG3209 189 VTLATGLAGSALLALGSGAILGGLAGAYSGSATTATGTALGTPASVAATVTGSATGAAGAGAAVATAATTLGGTTGAGTG 268
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1340 GNSGLLATK---SDETGWTTFFDYDSEGRLTNVTFPTGVVTNLHGDMDKAITVDIESSSREEDVSITSNLSSIDSFYTMV 1416
Cdd:COG3209 269 ASGAGLDAStgtGGAGGSNAAATAGGLGGAGLGSGGAGGGGTAGGTTTAAGTTGTAAVSGAADAGTTTTTGTGTGGTTTT 348
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1417 QDQLRNSYQIGYDGSLRIIYASGLDSHYQTEPHVLAGTANPTVAKRNMTLPGENGQNLVEWRFRKEQAQGKVNVFGRKLR 1496
Cdd:COG3209 349 VGGGGSLTLGGYGAAGGLTTSVGAGGGGSTSGSTTTVGGGGTATGSGGGSSTTGVGAGTTTTSTTGGDGGPATAAGALTA 428
|
330 340 350 360 370 380 390 400
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1497 VNGRNLLSVDFDRTTKTEKIYDDHRKFLLRIAYDTSGHPTLWLPSSKLMAVNVTYSSTGQIASIQRGTTSEKVDYDGQGR 1576
Cdd:COG3209 429 GGTATGTGTGGGGTTAGTDATTTTGGAGASGTLTTTGGAATGATTGGGTEAGTGGGTLTSGSAGATTLGTDTTLDDTLGG 508
|
410 420 430 440 450 460 470 480
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1577 IVSRVFADGKTWSYTYLEKSMVLLLHSQRQYIFEYDMWDRLSAITMPSVARHTMQTIRSIGYYRNIYNPPESNASIITDY 1656
Cdd:COG3209 509 TTTTTAGARGLVVTTGTTLTLGTTTTATLSATDATGTGDTTTTGTVGTGTSTGTGGTGTVTTTGDGTGGASTTTGTTGGT 588
|
490 500 510 520 530 540 550 560
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1657 NEEGLLLQTAFLGTSRRVLFKYRRQTRLSEILYDSTRVSFTYDETAGVLKTVNLQSDGFicTIRYRQIGPLIDRQIFRFS 1736
Cdd:COG3209 589 ATTTTVTTTTTTSTAGTTTTTTSGYTRAGLTLTLGTGTASGLERATASTGSTTGGTTGT--GVTTTGTTTTRATGTTGTG 666
|
570 580 590 600 610 620 630 640
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1737 EDGMVNARFDYSYDNSFRVTSMQGVINETPLPIDLYQFDDISGKVEQFGKFGVIYYDINQIISTAVMTYTKHFDAHGRIK 1816
Cdd:COG3209 667 TGVTAGLTTLATGGTTVGGGTGTTSTATTGATTGGTETGTTVTTLAGGTTTRLGTTTTGGGGGTTTDGTGTGGTTGTLTT 746
|
650 660 670 680 690 700 710 720
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1817 EIQYEifRSLMYWITIQYDNMGRVTKREIKIGPFANTTKYAYEYDVDGQLQTVYLNEKIMWRYNYDLNGNLHLLNPSNSA 1896
Cdd:COG3209 747 TSTTT--TTTAGALTYTYDALGRLTSETTPGGVTQGTYTTRYTYDALGRLTSVTYPDGETVTYTYDALGRLTSVITVGSG 824
|
730 740 750 760 770 780 790 800
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1897 RLTPL-----RYDLRDRITRLgdvqyrldEDGFLRQRGTEIFEYSSKGLLTRVYSKGSGWTviYRYDGLGRRVSSKTSLG 1971
Cdd:COG3209 825 GGTDLqdrtyTYDAAGNITSI--------TDALRAGTLTQTYTYDALGRLTSATDPGTTES--YTYDANGNLTSRTDGGT 894
|
810 820 830 840 850 860 870 880
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1972 QHLQFFYADLtyPTRITHvynhSSSEITSLYYDLQGHlfameissgdefyiaSDNTGTPLAVFSSNGLMLKQIQYTAYGE 2051
Cdd:COG3209 895 TTYTYDALGR--LVSVTK----PDGTTTTYTYDALGH---------------TDHLGSVRALTDASGQVVWRYDYDPFGN 953
|
890 900 910 920 930 940
....*....|....*....|....*....|....*....|....*....|....*....|....*....
gi 2421800221 2052 IYFDSNIDFQLVIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDPAPfNLYMFRNNNP 2120
Cdd:COG3209 954 LLAETSGAAANPLRFTGQEYDAETGLYYNGARYYDPALGRFLSPD-----PIGLAGGL-NLYAYVGNNP 1016
|
|
| Ten_N |
pfam06484 |
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of ... |
1-36 |
2.67e-18 |
|
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of the Teneurin family of proteins. These proteins are 'pair-rule' genes and are involved in tissue patterning, specifically probably neural patterning. The intracellular domain is cleaved in response to homophilic interaction of the extracellular domain, and translocates to the nucleus. Here it probably carries out to some transcriptional regulatory activity. The length of this region and the conservation suggests that there may be two structural domains here (personal obs:C Yeats).
Pssm-ID: 461932 [Multi-domain] Cd Length: 367 Bit Score: 89.26 E-value: 2.67e-18
10 20 30
....*....|....*....|....*....|....*.
gi 2421800221 1 MASGSVYSPPTRPLPRNTLSRSAFKFKKSSKYCSWK 36
Cdd:pfam06484 332 LTSGTVYSPPPRPLPRNTFSRPAFKLKKPYKYCSWK 367
|
|
| Vgb |
COG4257 |
Streptogramin lyase [Defense mechanisms]; |
896-1174 |
2.84e-10 |
|
Streptogramin lyase [Defense mechanisms];
Pssm-ID: 443399 [Multi-domain] Cd Length: 270 Bit Score: 63.50 E-value: 2.84e-10
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 896 DGNKLLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELSSNPAHRYYLATDPvTGDLYVSDTNTRRIYRpksLT 973
Cdd:COG4257 54 PLGGGSGPHGIAVDPDGNLWFTDNgnNRIGRIDPKTGEITTFALPGGGSNPHGIAFDP-DGNLWFTDQGGNRIGR---LD 129
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 974 gakdlTKNAEVVAGTgeqcLPFDEARcgdggkaveatlmsPKGMAVDKNGLIYFVD--GTMIRKVD-QNGIISTLLGSND 1050
Cdd:COG4257 130 -----PATGEVTEFP----LPTGGAG--------------PYGIAVDPDGNLWVTDfgANAIGRIDpDTGTLTEYALPTP 186
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1051 LTSarpltcdtsmhisqvrlewPTDLAINPmDNSIYVLD--NNVVLQITEnrqvriAAGRpmhcqvpgveypVGKHAVQT 1128
Cdd:COG4257 187 GAG-------------------PRGLAVDP-DGNLWVADtgSGRIGRFDP------KTGT------------VTEYPLPG 228
|
250 260 270 280
....*....|....*....|....*....|....*....|....*.
gi 2421800221 1129 TLESATAIAVSYSGVLYITETDekkINRIRQVTTDGEISLVAgIPS 1174
Cdd:COG4257 229 GGARPYGVAVDGDGRVWFAESG---ANRIVRFDPDTELTEYV-LPS 270
|
|
| Rhs_assc_core |
TIGR03696 |
RHS repeat-associated core domain; This model represents a conserved unique core sequence ... |
2046-2120 |
4.35e-09 |
|
RHS repeat-associated core domain; This model represents a conserved unique core sequence shared by large numbers of proteins. It is occasional in the Archaea Methanosarcina barkeri) but common in bacteria and eukaryotes. Most fall into two large classes. One class consists of long proteins in which two classes of repeats are abundant: an FG-GAP repeat (pfam01839) class, and an RHS repeat (pfam05593) or YD repeat (TIGR01643). This class includes secreted bacterial insecticidal toxins and intercellular signalling proteins such as the teneurins in animals. The other class consists of uncharacterized proteins shorter than 400 amino acids, where this core domain of about 75 amino acids tends to occur in the N-terminal half. Over twenty such proteins are found in Pseudomonas putida alone; little sequence similarity or repeat structure is found among these proteins outside the region modeled by this domain.
Pssm-ID: 274730 [Multi-domain] Cd Length: 77 Bit Score: 55.20 E-value: 4.35e-09
10 20 30 40 50 60 70
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*
gi 2421800221 2046 YTAYGEIYFDSNIDFQLvIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDpAPFNLYMFRNNNP 2120
Cdd:TIGR03696 1 YDPYGEVLSESGAAPNP-LRFTGQYYDAETGLYYNGARYYDPELGRFLSPD-----PIGLG-GGLNLYAYVGNNP 68
|
|
| acid_disulf_rpt |
NF033662 |
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with ... |
504-534 |
1.05e-07 |
|
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with four nearly invariant Cys residues in a repeat length of about 35 amino acids.
Pssm-ID: 411265 [Multi-domain] Cd Length: 32 Bit Score: 49.82 E-value: 1.05e-07
10 20 30
....*....|....*....|....*....|.
gi 2421800221 504 AMETLCTDSKDNEGDGLIDCMDPDCCLQSSC 534
Cdd:NF033662 2 ATDTTCSDGIDNDGDGLTDCADPDCAGNPVC 32
|
|
| DUF5885 |
pfam19232 |
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown ... |
250-404 |
1.17e-07 |
|
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown function found in viruses.
Pssm-ID: 437064 Cd Length: 265 Bit Score: 55.40 E-value: 1.17e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 250 CHGNGECvsGTCH-CFPGFLGPDCSraACPVLCSGNGQ----------YSKGRC----LCFSGwkgTECDVPTTQCI-DP 313
Cdd:pfam19232 34 CTTDAQC--GTCMtCVAGACTPKAS--CCGGVTCGAGQtcdaktntcvYVKGYCsadhPCPSG---SACDTAKNACIaQP 106
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 314 QCG---GRGiCIMG-------------------SCACNSGYK-GESCE--------EADCIDP---------------GC 347
Cdd:pfam19232 107 PYGpdsGKG-CVRGfgawiweldpatnsgvwrcRCANGSLYNsAHECSpladqtlcAAENLDPnalvpassvpafaayGW 185
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 348 SNHGVCIH-------------GECHCSPGWGGSNCEILKTmcpdqCSGHGTYLQESGSCTCD------------PN---W 399
Cdd:pfam19232 186 GNQPVLINkstagaavpsplaGVCPCKPGWAGGSCTEDRT-----CNGRGTWNETTGQCACNidfsghnscgddNNctsW 260
|
....*
gi 2421800221 400 TGPDC 404
Cdd:pfam19232 261 TGPRC 265
|
|
| C_rich_MXAN6577 |
NF041328 |
MXAN_6577-like cysteine-rich domain; |
301-455 |
3.06e-07 |
|
MXAN_6577-like cysteine-rich domain;
Pssm-ID: 469225 [Multi-domain] Cd Length: 145 Bit Score: 52.07 E-value: 3.06e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 301 TECDVPTTQCIDPQ--CGGRGICIMGScACNSGykgeSCEEAdcidpgCSNHGVCIHGECHCSPGwggsnceilKTMCPD 378
Cdd:NF041328 12 AGCPEPGAVCPEGLsvCGGACVDLRSD-PSNCG----ACGVA------CGAGQTCVAGACGCGPG---------TVACGG 71
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 379 QCSGHGTylqesgsctcDPNWTGPdcsneiCSVDCGSHGVCMGGTCR--CEEGWT--GPAC--------NQRACHPRCAE 446
Cdd:NF041328 72 ACVDTAS----------DPAHCGA------CGAACAPGQVCEGGACReaCSEGLTrcGGACvdlatdplHCGACGVACDP 135
|
....*....
gi 2421800221 447 HGTCKDGKC 455
Cdd:NF041328 136 GESCRGGAC 144
|
|
| C_rich_MXAN6577 |
NF041328 |
MXAN_6577-like cysteine-rich domain; |
409-495 |
4.00e-06 |
|
MXAN_6577-like cysteine-rich domain;
Pssm-ID: 469225 [Multi-domain] Cd Length: 145 Bit Score: 48.60 E-value: 4.00e-06
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 409 CSVDCGSHGVCMGGTCRCEEGWT--GPAC--------NQRACHPRCAEHGTCKDGKCEcsqgwngehctiEGCP-GLCNS 477
Cdd:NF041328 45 CGVACGAGQTCVAGACGCGPGTVacGGACvdtasdpaHCGACGAACAPGQVCEGGACR------------EACSeGLTRC 112
|
90 100
....*....|....*....|....
gi 2421800221 478 NGRCT-LDQNGWHC-----VCQPG 495
Cdd:NF041328 113 GGACVdLATDPLHCgacgvACDPG 136
|
|
| PLN02919 |
PLN02919 |
haloacid dehalogenase-like hydrolase family protein |
947-1230 |
7.34e-06 |
|
haloacid dehalogenase-like hydrolase family protein
Pssm-ID: 215497 [Multi-domain] Cd Length: 1057 Bit Score: 51.78 E-value: 7.34e-06
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 947 LATDPVTGDLYVSDTNTRRIYrpksltgAKDLTKNAEV-VAGTGEQCL---PFDEArcgdggkaveaTLMSPKGMAVD-K 1021
Cdd:PLN02919 573 LAIDLLNNRLFISDSNHNRIV-------VTDLDGNFIVqIGSTGEEGLrdgSFEDA-----------TFNRPQGLAYNaK 634
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1022 NGLIYFVD--GTMIRKVD-QNGIISTLLGS----NDLTSARPLTcdtsmhiSQVrLEWPTDLAINPMDNSIYV------- 1087
Cdd:PLN02919 635 KNLLYVADteNHALREIDfVNETVRTLAGNgtkgSDYQGGKKGT-------SQV-LNSPWDVCFEPVNEKVYIamagqhq 706
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1088 ------LD-------------------------------------NNVVLQITENRQVR-----------IAAGRPMhcq 1113
Cdd:PLN02919 707 iweyniSDgvtrvfsgdgyernlngssgtstsfaqpsgislspdlKELYIADSESSSIRaldlktggsrlLAGGDPT--- 783
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1114 VPGVEYPVGKH---AVQTTLESATAIAVSYSGVLYITETDEKKINRIRQVTtdGEISLVAGIPsecdckndancdcyQSG 1190
Cdd:PLN02919 784 FSDNLFKFGDHdgvGSEVLLQHPLGVLCAKDGQIYVADSYNHKIKKLDPAT--KRVTTLAGTG--------------KAG 847
|
330 340 350 360
....*....|....*....|....*....|....*....|..
gi 2421800221 1191 --DGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKNK 1230
Cdd:PLN02919 848 fkDGKALKAQLSEPAGLALGENGRLFVADTNNSLIRYLDLNK 889
|
|
| DSL |
pfam01414 |
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ... |
426-468 |
5.66e-05 |
|
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.
Pssm-ID: 460202 Cd Length: 46 Bit Score: 42.61 E-value: 5.66e-05
10 20 30 40
....*....|....*....|....*....|....*....|....*.
gi 2421800221 426 CEEGWTGPACNqRACHPRCAE--HGTC-KDGKCECSQGWNGEHCTI 468
Cdd:pfam01414 1 CDENYYGSTCS-KFCRPRDDKfgHYTCdANGNKVCLPGWTGPYCDK 45
|
|
| NHL |
cd05819 |
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ... |
1195-1321 |
1.56e-04 |
|
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.
Pssm-ID: 271320 [Multi-domain] Cd Length: 269 Bit Score: 46.16 E-value: 1.56e-04
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1195 KDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKN-KPLLN-------SMNFYE---VASPTDQELYI----------FD 1253
Cdd:cd05819 3 GPGELNNPQGIAVDSSGNIYVADTGNNRIQVFDPDgNFITSfgsfgsgDGQFNEpagVAVDSDGNLYVadtgnhriqkFD 82
|
90 100 110 120 130 140 150
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....
gi 2421800221 1254 INGTHQYTVSlVTGDYLYNFSY------SNDNDItAVTDSNGNtlRIrrdpnrmpvRVVSPDNQVIwLTIGTNG 1321
Cdd:cd05819 83 PDGNFLASFG-GSGDGDGEFNGprgiavDSSGNI-YVADTGNH--RI---------QKFDPDGEFL-TTFGSGG 142
|
|
| EGF_CA |
cd00054 |
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ... |
472-501 |
3.88e-03 |
|
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Pssm-ID: 238011 Cd Length: 38 Bit Score: 36.85 E-value: 3.88e-03
10 20 30
....*....|....*....|....*....|
gi 2421800221 472 PGLCNSNGRCTLDQNGWHCVCQPGWRGAGC 501
Cdd:cd00054 8 GNPCQNGGTCVNTVGSYRCSCPPGYTGRNC 37
|
|
| RHS_repeat |
pfam05593 |
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be ... |
1343-1374 |
7.47e-03 |
|
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be involved in ligand binding. Note that this model may not find all the repeats in a protein and that it covers two RHS repeats. The 3D structure of an RHS-repeat-containing protein (the B and C components of an ABC toxin complex) has been determined. The RHS repeats form an extended strip of beta-sheet that spirals around to form a hollow shell, encapsulating the variable C-terminal domain.
Pssm-ID: 461685 [Multi-domain] Cd Length: 37 Bit Score: 36.04 E-value: 7.47e-03
10 20 30
....*....|....*....|....*....|..
gi 2421800221 1343 GLLATKSDETGWTTFFDYDSEGRLTNVTFPTG 1374
Cdd:pfam05593 5 GRLTSVTDPDGRVTTYTYDAAGRLTAVTDPDG 36
|
|
|
|
Name |
Accession |
Description |
Interval |
E-value |
| NHL_like_1 |
cd14953 |
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ... |
874-1226 |
1.81e-40 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271323 [Multi-domain] Cd Length: 323 Bit Score: 153.45 E-value: 1.81e-40
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 874 VSSIMGNGRRrsiscpscnGQADGNKLLA----PVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVL----ELSSN--- 940
Cdd:cd14953 1 VSTVAGSGTA---------GFSGGGGTAArfnsPSGVAVDAAGNLYVADRgnHRIRKITPDGVVTTVAgtgtAGFADggg 71
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 941 PAHRYY----LATDPvTGDLYVSDTNTRRIYRpksLTGAKDLTknaeVVAGTGEqclpfdeARCGDGGKAVEATLMSPKG 1016
Cdd:cd14953 72 AAAQFNtpsgVAVDA-AGNLYVADTGNHRIRK---ITPDGVVS----TLAGTGT-------AGFSDDGGATAAQFNYPTG 136
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1017 MAVDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLTSAR--PLTcdtsmhisQVRLEWPTDLAINPMDNsIYVLD--N 1090
Cdd:cd14953 137 VAVDAAGNLYVADTGnhRIRKITPDGVVTTVAGTGGAGYAGdgPAT--------AAQFNNPTGVAVDAAGN-LYVADrgN 207
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1091 NVVLQITENRQVRIAAGRPmhcqvpGVEYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVA 1170
Cdd:cd14953 208 HRIRKITPDGVVTTVAGTG------TAGFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGN---HRIRKITPAGVVTTVA 278
|
330 340 350 360 370
....*....|....*....|....*....|....*....|....*....|....*..
gi 2421800221 1171 GIPSEcdckndancdcyQSGD-GYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAV 1226
Cdd:cd14953 279 GGGAG------------FSGDgGPATSAQFNNPTGVAVDAAGNLYVADTGNNRIRKI 323
|
|
| Tox-GHH |
pfam15636 |
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH ... |
2343-2420 |
6.94e-40 |
|
GHH signature containing HNH/Endo VII superfamily nuclease toxin; A predicted toxin of the HNH/Endonuclease VII fold present in bacterial polymorphic toxin systems with a characteriztic sG[HQ]H signature motif. In bacterial polymorphic toxin systems, the toxin is exported by the type 2, type 6, type 7 or TcdB/TcaC-type secretion system. The metazoan teneurin proteins possess an inactive of this domain at their C-terminus.
Pssm-ID: 464783 Cd Length: 78 Bit Score: 142.75 E-value: 6.94e-40
10 20 30 40 50 60 70
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*...
gi 2421800221 2343 EEKARILEQARQRALARAWAREQQRVRDGEEGARLWTEGEKRQLLSAGKVQGYDGYYVLSVEQYPELADSANNIQFLR 2420
Cdd:pfam15636 1 EERKRLLEHAKKRAVREAWHRERQLLRNGLPGSRDWTDEEKEELLSTGSVPGYDGEYIHPVEQYPELADDPSNIRFRK 78
|
|
| NHL_like_1 |
cd14953 |
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ... |
947-1227 |
4.70e-38 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271323 [Multi-domain] Cd Length: 323 Bit Score: 146.52 E-value: 4.70e-38
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 947 LATDPvTGDLYVSDTNTRRIYRpksltgakdLTKNAEV--VAGTGEqclpfdEARCGDGGKAveATLMSPKGMAVDKNGL 1024
Cdd:cd14953 28 VAVDA-AGNLYVADRGNHRIRK---------ITPDGVVttVAGTGT------AGFADGGGAA--AQFNTPSGVAVDAAGN 89
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1025 IYFVDGT--MIRKVDQNGIISTLLGsndlTSARPLTCDTSMhiSQVRLEWPTDLAINPMDNsIYVLD--NNVVLQITENR 1100
Cdd:cd14953 90 LYVADTGnhRIRKITPDGVVSTLAG----TGTAGFSDDGGA--TAAQFNYPTGVAVDAAGN-LYVADtgNHRIRKITPDG 162
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1101 QVRIAAGRPmhcqVPGveYPVGKHAVQTTLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVAGIPSEcdckn 1180
Cdd:cd14953 163 VVTTVAGTG----GAG--YAGDGPATAAQFNNPTGVAVDAAGNLYVADRGN---HRIRKITPDGVVTTVAGTGTA----- 228
|
250 260 270 280
....*....|....*....|....*....|....*....|....*..
gi 2421800221 1181 dancdcYQSGDGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVS 1227
Cdd:cd14953 229 ------GFSGDGGATAAQLNNPTGVAVDAAGNLYVADSGNHRIRKIT 269
|
|
| RhsA |
COG3209 |
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction ... |
1180-2120 |
1.76e-33 |
|
Uncharacterized conserved protein RhaS, contains 28 RHS repeats [General function prediction only];
Pssm-ID: 442442 [Multi-domain] Cd Length: 1103 Bit Score: 142.20 E-value: 1.76e-33
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1180 NDANCDCYQSGDGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKNKPLLNSMNFYEVASPTDQELYIFDINGTHQ 1259
Cdd:COG3209 109 AAATASAGRLVSTGAGAGGTVTAATGGTLGATAGSATTGSTDGGRGGVAVTGLAGGGASAYGLTLGGAAAGPATGVGTGA 188
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1260 YTVSLVTGDYLYNFSYSNDNDITAVTDSNGNTLRIRRDPNRMPVRVVSPDNQVIWLTIGTNGCLKSMTAQGLELVLFTYH 1339
Cdd:COG3209 189 VTLATGLAGSALLALGSGAILGGLAGAYSGSATTATGTALGTPASVAATVTGSATGAAGAGAAVATAATTLGGTTGAGTG 268
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1340 GNSGLLATK---SDETGWTTFFDYDSEGRLTNVTFPTGVVTNLHGDMDKAITVDIESSSREEDVSITSNLSSIDSFYTMV 1416
Cdd:COG3209 269 ASGAGLDAStgtGGAGGSNAAATAGGLGGAGLGSGGAGGGGTAGGTTTAAGTTGTAAVSGAADAGTTTTTGTGTGGTTTT 348
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1417 QDQLRNSYQIGYDGSLRIIYASGLDSHYQTEPHVLAGTANPTVAKRNMTLPGENGQNLVEWRFRKEQAQGKVNVFGRKLR 1496
Cdd:COG3209 349 VGGGGSLTLGGYGAAGGLTTSVGAGGGGSTSGSTTTVGGGGTATGSGGGSSTTGVGAGTTTTSTTGGDGGPATAAGALTA 428
|
330 340 350 360 370 380 390 400
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1497 VNGRNLLSVDFDRTTKTEKIYDDHRKFLLRIAYDTSGHPTLWLPSSKLMAVNVTYSSTGQIASIQRGTTSEKVDYDGQGR 1576
Cdd:COG3209 429 GGTATGTGTGGGGTTAGTDATTTTGGAGASGTLTTTGGAATGATTGGGTEAGTGGGTLTSGSAGATTLGTDTTLDDTLGG 508
|
410 420 430 440 450 460 470 480
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1577 IVSRVFADGKTWSYTYLEKSMVLLLHSQRQYIFEYDMWDRLSAITMPSVARHTMQTIRSIGYYRNIYNPPESNASIITDY 1656
Cdd:COG3209 509 TTTTTAGARGLVVTTGTTLTLGTTTTATLSATDATGTGDTTTTGTVGTGTSTGTGGTGTVTTTGDGTGGASTTTGTTGGT 588
|
490 500 510 520 530 540 550 560
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1657 NEEGLLLQTAFLGTSRRVLFKYRRQTRLSEILYDSTRVSFTYDETAGVLKTVNLQSDGFicTIRYRQIGPLIDRQIFRFS 1736
Cdd:COG3209 589 ATTTTVTTTTTTSTAGTTTTTTSGYTRAGLTLTLGTGTASGLERATASTGSTTGGTTGT--GVTTTGTTTTRATGTTGTG 666
|
570 580 590 600 610 620 630 640
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1737 EDGMVNARFDYSYDNSFRVTSMQGVINETPLPIDLYQFDDISGKVEQFGKFGVIYYDINQIISTAVMTYTKHFDAHGRIK 1816
Cdd:COG3209 667 TGVTAGLTTLATGGTTVGGGTGTTSTATTGATTGGTETGTTVTTLAGGTTTRLGTTTTGGGGGTTTDGTGTGGTTGTLTT 746
|
650 660 670 680 690 700 710 720
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1817 EIQYEifRSLMYWITIQYDNMGRVTKREIKIGPFANTTKYAYEYDVDGQLQTVYLNEKIMWRYNYDLNGNLHLLNPSNSA 1896
Cdd:COG3209 747 TSTTT--TTTAGALTYTYDALGRLTSETTPGGVTQGTYTTRYTYDALGRLTSVTYPDGETVTYTYDALGRLTSVITVGSG 824
|
730 740 750 760 770 780 790 800
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1897 RLTPL-----RYDLRDRITRLgdvqyrldEDGFLRQRGTEIFEYSSKGLLTRVYSKGSGWTviYRYDGLGRRVSSKTSLG 1971
Cdd:COG3209 825 GGTDLqdrtyTYDAAGNITSI--------TDALRAGTLTQTYTYDALGRLTSATDPGTTES--YTYDANGNLTSRTDGGT 894
|
810 820 830 840 850 860 870 880
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1972 QHLQFFYADLtyPTRITHvynhSSSEITSLYYDLQGHlfameissgdefyiaSDNTGTPLAVFSSNGLMLKQIQYTAYGE 2051
Cdd:COG3209 895 TTYTYDALGR--LVSVTK----PDGTTTTYTYDALGH---------------TDHLGSVRALTDASGQVVWRYDYDPFGN 953
|
890 900 910 920 930 940
....*....|....*....|....*....|....*....|....*....|....*....|....*....
gi 2421800221 2052 IYFDSNIDFQLVIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDPAPfNLYMFRNNNP 2120
Cdd:COG3209 954 LLAETSGAAANPLRFTGQEYDAETGLYYNGARYYDPALGRFLSPD-----PIGLAGGL-NLYAYVGNNP 1016
|
|
| NHL_like_1 |
cd14953 |
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ... |
984-1227 |
3.35e-31 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271323 [Multi-domain] Cd Length: 323 Bit Score: 126.49 E-value: 3.35e-31
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 984 VVAGTGeqclpfdeARCGDGGKAVEATLMSPKGMAVDKNGLIYFVDGT--MIRKVDQNGIISTLLG------SNDLTSAr 1055
Cdd:cd14953 3 TVAGSG--------TAGFSGGGGTAARFNSPSGVAVDAAGNLYVADRGnhRIRKITPDGVVTTVAGtgtagfADGGGAA- 73
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1056 pltcdtsmhisqVRLEWPTDLAINPMDNsIYVLD--NNVVLQITENRQVRIAAGrpmhcqVPGVEYPVGKHAVQTTLESA 1133
Cdd:cd14953 74 ------------AQFNTPSGVAVDAAGN-LYVADtgNHRIRKITPDGVVSTLAG------TGTAGFSDDGGATAAQFNYP 134
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1134 TAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISLVAGIPSEcdckndancdcYQSGDGYAKDAKLSAPSSLAASPDGTL 1213
Cdd:cd14953 135 TGVAVDAAGNLYVADTGN---HRIRKITPDGVVTTVAGTGGA-----------GYAGDGPATAAQFNNPTGVAVDAAGNL 200
|
250
....*....|....
gi 2421800221 1214 YIADLGNIRIRAVS 1227
Cdd:cd14953 201 YVADRGNHRIRKIT 214
|
|
| NHL |
cd05819 |
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ... |
893-1224 |
2.30e-18 |
|
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.
Pssm-ID: 271320 [Multi-domain] Cd Length: 269 Bit Score: 87.37 E-value: 2.30e-18
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 893 GQADGnKLLAPVALACGIDGSLYVGDF--NYVRRIFPSGN-VTSVLELSSNPAHRYY---LATDPvTGDLYVSDTNTRRI 966
Cdd:cd05819 1 GTGPG-ELNNPQGIAVDSSGNIYVADTgnNRIQVFDPDGNfITSFGSFGSGDGQFNEpagVAVDS-DGNLYVADTGNHRI 78
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 967 YRpksltgakdLTKNAEVVAGTGeqclpfdearcGDGGKAVEatLMSPKGMAVDKNGLIYFVDgTM---IRKVDQNGIIS 1043
Cdd:cd05819 79 QK---------FDPDGNFLASFG-----------GSGDGDGE--FNGPRGIAVDSSGNIYVAD-TGnhrIQKFDPDGEFL 135
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1044 TLLGSNDLTSARpltcdtsmhisqvrLEWPTDLAINPmDNSIYVLDnnvvlqiTENRQVRI--AAGRPMHcQVPGVEYPV 1121
Cdd:cd05819 136 TTFGSGGSGPGQ--------------FNGPTGVAVDS-DGNIYVAD-------TGNHRIQVfdPDGNFLT-TFGSTGTGP 192
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1122 GKhavqttLESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEISlvagipsecdckndancdcYQSGDGYAKDAKLSA 1201
Cdd:cd05819 193 GQ------FNYPTGIAVDSDGNIYVADSGN---NRVQVFDPDGAGF-------------------GGNGNFLGSDGQFNR 244
|
330 340
....*....|....*....|...
gi 2421800221 1202 PSSLAASPDGTLYIADLGNIRIR 1224
Cdd:cd05819 245 PSGLAVDSDGNLYVADTGNNRIQ 267
|
|
| Ten_N |
pfam06484 |
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of ... |
1-36 |
2.67e-18 |
|
Teneurin Intracellular Region; This family is found in the intracellular N-terminal region of the Teneurin family of proteins. These proteins are 'pair-rule' genes and are involved in tissue patterning, specifically probably neural patterning. The intracellular domain is cleaved in response to homophilic interaction of the extracellular domain, and translocates to the nucleus. Here it probably carries out to some transcriptional regulatory activity. The length of this region and the conservation suggests that there may be two structural domains here (personal obs:C Yeats).
Pssm-ID: 461932 [Multi-domain] Cd Length: 367 Bit Score: 89.26 E-value: 2.67e-18
10 20 30
....*....|....*....|....*....|....*.
gi 2421800221 1 MASGSVYSPPTRPLPRNTLSRSAFKFKKSSKYCSWK 36
Cdd:pfam06484 332 LTSGTVYSPPPRPLPRNTFSRPAFKLKKPYKYCSWK 367
|
|
| NHL |
cd05819 |
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ... |
892-1157 |
1.42e-15 |
|
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.
Pssm-ID: 271320 [Multi-domain] Cd Length: 269 Bit Score: 79.28 E-value: 1.42e-15
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 892 NGQADGNkLLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELSSNPAHRYY----LATDPvTGDLYVSDTNTRR 965
Cdd:cd05819 47 FGSGDGQ-FNEPAGVAVDSDGNLYVADTgnHRIQKFDPDGNFLASFGGSGDGDGEFNgprgIAVDS-SGNIYVADTGNHR 124
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 966 IYRpksltgakdLTKNAEVVAGTGeqclpfdearcgdGGKAVEATLMSPKGMAVDKNGLIYFVDGT--MIRKVDQNGIIS 1043
Cdd:cd05819 125 IQK---------FDPDGEFLTTFG-------------SGGSGPGQFNGPTGVAVDSDGNIYVADTGnhRIQVFDPDGNFL 182
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1044 TLLGSNDLTSARpltcdtsmhisqvrLEWPTDLAINPMDNsIYVLD--NNVVLQITENRQVRIAAGRPMhCQVPGVEYPV 1121
Cdd:cd05819 183 TTFGSTGTGPGQ--------------FNYPTGIAVDSDGN-IYVADsgNNRVQVFDPDGAGFGGNGNFL-GSDGQFNRPS 246
|
250 260 270
....*....|....*....|....*....|....*.
gi 2421800221 1122 GkhavqttlesataIAVSYSGVLYITETDEKKINRI 1157
Cdd:cd05819 247 G-------------LAVDSDGNLYVADTGNNRIQVF 269
|
|
| Vgb |
COG4257 |
Streptogramin lyase [Defense mechanisms]; |
896-1174 |
2.84e-10 |
|
Streptogramin lyase [Defense mechanisms];
Pssm-ID: 443399 [Multi-domain] Cd Length: 270 Bit Score: 63.50 E-value: 2.84e-10
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 896 DGNKLLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELSSNPAHRYYLATDPvTGDLYVSDTNTRRIYRpksLT 973
Cdd:COG4257 54 PLGGGSGPHGIAVDPDGNLWFTDNgnNRIGRIDPKTGEITTFALPGGGSNPHGIAFDP-DGNLWFTDQGGNRIGR---LD 129
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 974 gakdlTKNAEVVAGTgeqcLPFDEARcgdggkaveatlmsPKGMAVDKNGLIYFVD--GTMIRKVD-QNGIISTLLGSND 1050
Cdd:COG4257 130 -----PATGEVTEFP----LPTGGAG--------------PYGIAVDPDGNLWVTDfgANAIGRIDpDTGTLTEYALPTP 186
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1051 LTSarpltcdtsmhisqvrlewPTDLAINPmDNSIYVLD--NNVVLQITEnrqvriAAGRpmhcqvpgveypVGKHAVQT 1128
Cdd:COG4257 187 GAG-------------------PRGLAVDP-DGNLWVADtgSGRIGRFDP------KTGT------------VTEYPLPG 228
|
250 260 270 280
....*....|....*....|....*....|....*....|....*.
gi 2421800221 1129 TLESATAIAVSYSGVLYITETDekkINRIRQVTTDGEISLVAgIPS 1174
Cdd:COG4257 229 GGARPYGVAVDGDGRVWFAESG---ANRIVRFDPDTELTEYV-LPS 270
|
|
| NHL_PKND_like |
cd14952 |
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein ... |
947-1223 |
1.01e-09 |
|
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein with a cytosolic kinase domain and an extracellular sensor domain that contains NHL repeats. It plays a key role in the development of central nervous system tuberculosis, by mediating the invasion of host brain endothelia. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271322 [Multi-domain] Cd Length: 247 Bit Score: 61.45 E-value: 1.01e-09
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 947 LATDPvTGDLYVSDTNTRRIYRpksltgakdltknaeVVAGTGEQ-CLPFDEarcgdggkaveatLMSPKGMAVDKNGLI 1025
Cdd:cd14952 15 VAVDA-AGNVYVADSGNNRVLK---------------LAAGSTTQtVLPFTG-------------LYQPQGVAVDAAGTV 65
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1026 YFVDGtmirkvDQNGIISTLLGSNDLTsARPLTcdtsmhisqvRLEWPTDLAINPMDNsIYVLDNnvvlqiTENRQVRIA 1105
Cdd:cd14952 66 YVTDF------GNNRVLKLAAGSTTQT-VLPFT----------GLNDPTGVAVDAAGN-VYVADT------GNNRVLKLA 121
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1106 AGRPMHCQVPgveypvgkhavQTTLESATAIAVSYSGVLYITETDEkkiNRIRQvttdgeisLVAGipsecdckndANCD 1185
Cdd:cd14952 122 AGSNTQTVLP-----------FTGLSNPDGVAVDGAGNVYVTDTGN---NRVLK--------LAAG----------STTQ 169
|
250 260 270
....*....|....*....|....*....|....*...
gi 2421800221 1186 CYQSGDGyakdakLSAPSSLAASPDGTLYIADLGNIRI 1223
Cdd:cd14952 170 TVLPFTG------LNSPSGVAVDTAGNVYVTDHGNNRV 201
|
|
| Rhs_assc_core |
TIGR03696 |
RHS repeat-associated core domain; This model represents a conserved unique core sequence ... |
2046-2120 |
4.35e-09 |
|
RHS repeat-associated core domain; This model represents a conserved unique core sequence shared by large numbers of proteins. It is occasional in the Archaea Methanosarcina barkeri) but common in bacteria and eukaryotes. Most fall into two large classes. One class consists of long proteins in which two classes of repeats are abundant: an FG-GAP repeat (pfam01839) class, and an RHS repeat (pfam05593) or YD repeat (TIGR01643). This class includes secreted bacterial insecticidal toxins and intercellular signalling proteins such as the teneurins in animals. The other class consists of uncharacterized proteins shorter than 400 amino acids, where this core domain of about 75 amino acids tends to occur in the N-terminal half. Over twenty such proteins are found in Pseudomonas putida alone; little sequence similarity or repeat structure is found among these proteins outside the region modeled by this domain.
Pssm-ID: 274730 [Multi-domain] Cd Length: 77 Bit Score: 55.20 E-value: 4.35e-09
10 20 30 40 50 60 70
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*
gi 2421800221 2046 YTAYGEIYFDSNIDFQLvIGFHGGLYDPLTKLIHFGERDYDILAGRWTTPDieiwkRIGKDpAPFNLYMFRNNNP 2120
Cdd:TIGR03696 1 YDPYGEVLSESGAAPNP-LRFTGQYYDAETGLYYNGARYYDPELGRFLSPD-----PIGLG-GGLNLYAYVGNNP 68
|
|
| NHL_like_2 |
cd14957 |
Uncharacterized NHL-repeat domain in bacterial and archaeal proteins; The NHL (NCL-1, HT2A and ... |
1011-1321 |
8.64e-08 |
|
Uncharacterized NHL-repeat domain in bacterial and archaeal proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271327 [Multi-domain] Cd Length: 280 Bit Score: 56.12 E-value: 8.64e-08
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1011 LMSPKGMAVDKNGLIYFVD--GTMIRKVDQNGIISTLLGSNDltsarpltcdtsmhISQVRLEWPTDLAINPMDNsIYVL 1088
Cdd:cd14957 17 FNTPRGIAVDSAGNIYVADtgNNRIQVFTSSGVYSYSIGSGG--------------TGSGQFNSPYGIAVDSNGN-IYVA 81
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1089 DNNvvlqitENR-QVRIAAGrpmhcqvpGVEYPVGKHAVQTT-LESATAIAVSYSGVLYITETDEkkiNRIRQVTTDGEI 1166
Cdd:cd14957 82 DTD------NNRiQVFNSSG--------VYQYSIGTGGSGDGqFNGPYGIAVDSNGNIYVADTGN---HRIQVFTSSGTF 144
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1167 slvagipsecdckndancdCYQSGDGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRavsknkpllnsmnfyevasptd 1246
Cdd:cd14957 145 -------------------SYSIGSGGTGPGQFNGPQGIAVDSDGNIYVADTGNHRIQ---------------------- 183
|
250 260 270 280 290 300 310
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*.
gi 2421800221 1247 qelyIFDINGTHQYTV-SLVTGDYLynFSYSNDNDItavtDSNGNTLRIRRDPNRmpVRVVSPDNqVIWLTIGTNG 1321
Cdd:cd14957 184 ----VFTSSGTFQYTFgSSGSGPGQ--FSDPYGIAV----DSDGNIYVADTGNHR--IQVFTSSG-AYQYSIGTSG 246
|
|
| acid_disulf_rpt |
NF033662 |
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with ... |
504-534 |
1.05e-07 |
|
acidic double-disulfide repeat; The acidic double-disulfide repeat is an Asp-rich repeat with four nearly invariant Cys residues in a repeat length of about 35 amino acids.
Pssm-ID: 411265 [Multi-domain] Cd Length: 32 Bit Score: 49.82 E-value: 1.05e-07
10 20 30
....*....|....*....|....*....|.
gi 2421800221 504 AMETLCTDSKDNEGDGLIDCMDPDCCLQSSC 534
Cdd:NF033662 2 ATDTTCSDGIDNDGDGLTDCADPDCAGNPVC 32
|
|
| DUF5885 |
pfam19232 |
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown ... |
250-404 |
1.17e-07 |
|
Family of unknown function (DUF5885); This is a family of uncharacterized proteins of unknown function found in viruses.
Pssm-ID: 437064 Cd Length: 265 Bit Score: 55.40 E-value: 1.17e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 250 CHGNGECvsGTCH-CFPGFLGPDCSraACPVLCSGNGQ----------YSKGRC----LCFSGwkgTECDVPTTQCI-DP 313
Cdd:pfam19232 34 CTTDAQC--GTCMtCVAGACTPKAS--CCGGVTCGAGQtcdaktntcvYVKGYCsadhPCPSG---SACDTAKNACIaQP 106
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 314 QCG---GRGiCIMG-------------------SCACNSGYK-GESCE--------EADCIDP---------------GC 347
Cdd:pfam19232 107 PYGpdsGKG-CVRGfgawiweldpatnsgvwrcRCANGSLYNsAHECSpladqtlcAAENLDPnalvpassvpafaayGW 185
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 348 SNHGVCIH-------------GECHCSPGWGGSNCEILKTmcpdqCSGHGTYLQESGSCTCD------------PN---W 399
Cdd:pfam19232 186 GNQPVLINkstagaavpsplaGVCPCKPGWAGGSCTEDRT-----CNGRGTWNETTGQCACNidfsghnscgddNNctsW 260
|
....*
gi 2421800221 400 TGPDC 404
Cdd:pfam19232 261 TGPRC 265
|
|
| NHL_PKND_like |
cd14952 |
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein ... |
900-1154 |
1.69e-07 |
|
NHL repeat domain of the protein kinase PknD; PknD is a mycobacterial transmembrane protein with a cytosolic kinase domain and an extracellular sensor domain that contains NHL repeats. It plays a key role in the development of central nervous system tuberculosis, by mediating the invasion of host brain endothelia. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271322 [Multi-domain] Cd Length: 247 Bit Score: 54.91 E-value: 1.69e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 900 LLAPVALACGIDGSLYVGDF--NYVRRIFPSGNVTSVLELS--SNPAHryyLATDPVtGDLYVSDTNTRRIyrpksltga 975
Cdd:cd14952 51 LYQPQGVAVDAAGTVYVTDFgnNRVLKLAAGSTTQTVLPFTglNDPTG---VAVDAA-GNVYVADTGNNRV--------- 117
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 976 kdltknAEVVAGTGEQC-LPFdearcgdggkaveATLMSPKGMAVDKNGLIYFVDGtmirkvDQNGIISTLLGSNDLTsA 1054
Cdd:cd14952 118 ------LKLAAGSNTQTvLPF-------------TGLSNPDGVAVDGAGNVYVTDT------GNNRVLKLAAGSTTQT-V 171
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1055 RPLTCDTSmhisqvrlewPTDLAINPMDNsIYVLDNNvvlqitENRQVRIAAGRPMHCQVP--GVEYPVGkhavqttles 1132
Cdd:cd14952 172 LPFTGLNS----------PSGVAVDTAGN-VYVTDHG------NNRVLKLAAGSTTPTVLPftGLNGPLG---------- 224
|
250 260
....*....|....*....|..
gi 2421800221 1133 ataIAVSYSGVLYITETDEKKI 1154
Cdd:cd14952 225 ---VAVDAAGNVYVADRGNDRV 243
|
|
| C_rich_MXAN6577 |
NF041328 |
MXAN_6577-like cysteine-rich domain; |
301-455 |
3.06e-07 |
|
MXAN_6577-like cysteine-rich domain;
Pssm-ID: 469225 [Multi-domain] Cd Length: 145 Bit Score: 52.07 E-value: 3.06e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 301 TECDVPTTQCIDPQ--CGGRGICIMGScACNSGykgeSCEEAdcidpgCSNHGVCIHGECHCSPGwggsnceilKTMCPD 378
Cdd:NF041328 12 AGCPEPGAVCPEGLsvCGGACVDLRSD-PSNCG----ACGVA------CGAGQTCVAGACGCGPG---------TVACGG 71
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 379 QCSGHGTylqesgsctcDPNWTGPdcsneiCSVDCGSHGVCMGGTCR--CEEGWT--GPAC--------NQRACHPRCAE 446
Cdd:NF041328 72 ACVDTAS----------DPAHCGA------CGAACAPGQVCEGGACReaCSEGLTrcGGACvdlatdplHCGACGVACDP 135
|
....*....
gi 2421800221 447 HGTCKDGKC 455
Cdd:NF041328 136 GESCRGGAC 144
|
|
| Vgb |
COG4257 |
Streptogramin lyase [Defense mechanisms]; |
945-1224 |
9.93e-07 |
|
Streptogramin lyase [Defense mechanisms];
Pssm-ID: 443399 [Multi-domain] Cd Length: 270 Bit Score: 52.71 E-value: 9.93e-07
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 945 YYLATDPvTGDLYVSDTNTRRIYRpksltgakdltknaeVVAGTGEqclpFDEARCGDGGkaveatlmSPKGMAVDKNGL 1024
Cdd:COG4257 20 RDVAVDP-DGAVWFTDQGGGRIGR---------------LDPATGE----FTEYPLGGGS--------GPHGIAVDPDGN 71
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1025 IYFVDGT--MIRKVD-QNGIISTLLGSNDLTSarpltcdtsmhisqvrlewPTDLAINPmDNSIYVLD--NNVVLQIT-E 1098
Cdd:COG4257 72 LWFTDNGnnRIGRIDpKTGEITTFALPGGGSN-------------------PHGIAFDP-DGNLWFTDqgGNRIGRLDpA 131
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1099 NRQVRiaagrpmhcqvpgvEYPVGKHAVQTTlesatAIAVSYSGVLYITETdekKINRIRQVTTD-GEISLvagipsecd 1177
Cdd:COG4257 132 TGEVT--------------EFPLPTGGAGPY-----GIAVDPDGNLWVTDF---GANAIGRIDPDtGTLTE--------- 180
|
250 260 270 280
....*....|....*....|....*....|....*....|....*..
gi 2421800221 1178 ckndancdcyqsgdgYAKDAKLSAPSSLAASPDGTLYIADLGNIRIR 1224
Cdd:COG4257 181 ---------------YALPTPGAGPRGLAVDPDGNLWVADTGSGRIG 212
|
|
| C_rich_MXAN6577 |
NF041328 |
MXAN_6577-like cysteine-rich domain; |
409-495 |
4.00e-06 |
|
MXAN_6577-like cysteine-rich domain;
Pssm-ID: 469225 [Multi-domain] Cd Length: 145 Bit Score: 48.60 E-value: 4.00e-06
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 409 CSVDCGSHGVCMGGTCRCEEGWT--GPAC--------NQRACHPRCAEHGTCKDGKCEcsqgwngehctiEGCP-GLCNS 477
Cdd:NF041328 45 CGVACGAGQTCVAGACGCGPGTVacGGACvdtasdpaHCGACGAACAPGQVCEGGACR------------EACSeGLTRC 112
|
90 100
....*....|....*....|....
gi 2421800221 478 NGRCT-LDQNGWHC-----VCQPG 495
Cdd:NF041328 113 GGACVdLATDPLHCgacgvACDPG 136
|
|
| PLN02919 |
PLN02919 |
haloacid dehalogenase-like hydrolase family protein |
947-1230 |
7.34e-06 |
|
haloacid dehalogenase-like hydrolase family protein
Pssm-ID: 215497 [Multi-domain] Cd Length: 1057 Bit Score: 51.78 E-value: 7.34e-06
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 947 LATDPVTGDLYVSDTNTRRIYrpksltgAKDLTKNAEV-VAGTGEQCL---PFDEArcgdggkaveaTLMSPKGMAVD-K 1021
Cdd:PLN02919 573 LAIDLLNNRLFISDSNHNRIV-------VTDLDGNFIVqIGSTGEEGLrdgSFEDA-----------TFNRPQGLAYNaK 634
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1022 NGLIYFVD--GTMIRKVD-QNGIISTLLGS----NDLTSARPLTcdtsmhiSQVrLEWPTDLAINPMDNSIYV------- 1087
Cdd:PLN02919 635 KNLLYVADteNHALREIDfVNETVRTLAGNgtkgSDYQGGKKGT-------SQV-LNSPWDVCFEPVNEKVYIamagqhq 706
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1088 ------LD-------------------------------------NNVVLQITENRQVR-----------IAAGRPMhcq 1113
Cdd:PLN02919 707 iweyniSDgvtrvfsgdgyernlngssgtstsfaqpsgislspdlKELYIADSESSSIRaldlktggsrlLAGGDPT--- 783
|
250 260 270 280 290 300 310 320
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1114 VPGVEYPVGKH---AVQTTLESATAIAVSYSGVLYITETDEKKINRIRQVTtdGEISLVAGIPsecdckndancdcyQSG 1190
Cdd:PLN02919 784 FSDNLFKFGDHdgvGSEVLLQHPLGVLCAKDGQIYVADSYNHKIKKLDPAT--KRVTTLAGTG--------------KAG 847
|
330 340 350 360
....*....|....*....|....*....|....*....|..
gi 2421800221 1191 --DGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKNK 1230
Cdd:PLN02919 848 fkDGKALKAQLSEPAGLALGENGRLFVADTNNSLIRYLDLNK 889
|
|
| DSL |
pfam01414 |
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ... |
426-468 |
5.66e-05 |
|
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.
Pssm-ID: 460202 Cd Length: 46 Bit Score: 42.61 E-value: 5.66e-05
10 20 30 40
....*....|....*....|....*....|....*....|....*.
gi 2421800221 426 CEEGWTGPACNqRACHPRCAE--HGTC-KDGKCECSQGWNGEHCTI 468
Cdd:pfam01414 1 CDENYYGSTCS-KFCRPRDDKfgHYTCdANGNKVCLPGWTGPYCDK 45
|
|
| DSL |
pfam01414 |
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain ... |
395-437 |
1.54e-04 |
|
Delta serrate ligand; This family has been redefined to correspond to the EGF-like domain defined by structure.
Pssm-ID: 460202 Cd Length: 46 Bit Score: 41.07 E-value: 1.54e-04
10 20 30 40
....*....|....*....|....*....|....*....|....*.
gi 2421800221 395 CDPNWTGPDCSNEiCSV--DCGSHGVC-MGGTCRCEEGWTGPACNQ 437
Cdd:pfam01414 1 CDENYYGSTCSKF-CRPrdDKFGHYTCdANGNKVCLPGWTGPYCDK 45
|
|
| NHL |
cd05819 |
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in ... |
1195-1321 |
1.56e-04 |
|
NHL repeat unit of beta-propeller proteins; The NHL(NCL-1, HT2A and LIN-41)-repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures. The repeats have a catalytic activity in Peptidyl-glycine alpha-amidating monooxygenase; proteolysis has shown that the Peptidyl-alpha-hydroxyglycine alpha-amidating lyase (PAL) activity is localized to the repeats. Tripartite motif-containing protein 32 interacts with the activation domain of Tat. This interaction is mediated by the NHL repeats.
Pssm-ID: 271320 [Multi-domain] Cd Length: 269 Bit Score: 46.16 E-value: 1.56e-04
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1195 KDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKN-KPLLN-------SMNFYE---VASPTDQELYI----------FD 1253
Cdd:cd05819 3 GPGELNNPQGIAVDSSGNIYVADTGNNRIQVFDPDgNFITSfgsfgsgDGQFNEpagVAVDSDGNLYVadtgnhriqkFD 82
|
90 100 110 120 130 140 150
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....
gi 2421800221 1254 INGTHQYTVSlVTGDYLYNFSY------SNDNDItAVTDSNGNtlRIrrdpnrmpvRVVSPDNQVIwLTIGTNG 1321
Cdd:cd05819 83 PDGNFLASFG-GSGDGDGEFNGprgiavDSSGNI-YVADTGNH--RI---------QKFDPDGEFL-TTFGSGG 142
|
|
| NHL_like_1 |
cd14953 |
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat ... |
1189-1229 |
2.12e-04 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; This bacterial family of NHL-repeat domains is found in a variety of domain architectures. The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271323 [Multi-domain] Cd Length: 323 Bit Score: 45.98 E-value: 2.12e-04
10 20 30 40
....*....|....*....|....*....|....*....|.
gi 2421800221 1189 SGDGYAKDAKLSAPSSLAASPDGTLYIADLGNIRIRAVSKN 1229
Cdd:cd14953 12 FSGGGGTAARFNSPSGVAVDAAGNLYVADRGNHRIRKITPD 52
|
|
| Keratin_B2 |
pfam01500 |
Keratin, high sulfur B2 protein; High sulfur proteins are cysteine-rich proteins synthesized ... |
337-501 |
3.26e-04 |
|
Keratin, high sulfur B2 protein; High sulfur proteins are cysteine-rich proteins synthesized during the differentiation of hair matrix cells, and form hair fibres in association with hair keratin intermediate filaments. This family has been divided up into four regions, with the second region containing 8 copies of a short repeat. This family is also known as B2 or KAP1.
Pssm-ID: 366678 [Multi-domain] Cd Length: 161 Bit Score: 43.63 E-value: 3.26e-04
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 337 CEEADCIDPGCSNHGVCihGECHCSPGWGGSNCeILKTMCPDQCSGHGTYLQESGSCTCDPNwTGPDCSNEICSVDCGSH 416
Cdd:pfam01500 4 CGTSFCGFPTCSTGGTC--GSGCCQPCCCQSSC-CRPSCCQTSCCQPTTFQSSCCRPTCQPC-CQTSCCQPTCCQTSSCQ 79
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 417 GVCMGGTCRCEEGWTGPACNQRACHPRCAEHGTCKDGKCECSqgwngehCTIEGCPGLCNSNGRCtldqngwhcvCQPGW 496
Cdd:pfam01500 80 TGCGGIGYGQEGSSGAVSSRTRWCRPDCRVEGTCLPPCCVVS-------CTPPTCCQLHHAQASC----------CRPSY 142
|
....*
gi 2421800221 497 RGAGC 501
Cdd:pfam01500 143 CGQSC 147
|
|
| YvrE |
COG3386 |
Sugar lactone lactonase YvrE [Carbohydrate transport and metabolism]; Sugar lactone lactonase ... |
906-1042 |
3.53e-04 |
|
Sugar lactone lactonase YvrE [Carbohydrate transport and metabolism]; Sugar lactone lactonase YvrE is part of the Pathway/BioSystem: Non-phosphorylated Entner-Doudoroff pathway
Pssm-ID: 442613 [Multi-domain] Cd Length: 266 Bit Score: 44.88 E-value: 3.53e-04
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 906 LACGIDGSLYVGDFNYVR------RIFPSGNVTSVLE--LSSN-----PAHRYylatdpvtgdLYVSDTNTRRIYR-PKS 971
Cdd:COG3386 98 GVVDPDGRLYFTDMGEYLptgalyRVDPDGSLRVLADglTFPNgiafsPDGRT----------LYVADTGAGRIYRfDLD 167
|
90 100 110 120 130 140 150
....*....|....*....|....*....|....*....|....*....|....*....|....*....|...
gi 2421800221 972 LTGAkdLTkNAEVVAgtgeqclpfdEARCGDGGkaveatlmsPKGMAVDKNGLIY--FVDGTMIRKVDQNGII 1042
Cdd:COG3386 168 ADGT--LG-NRRVFA----------DLPDGPGG---------PDGLAVDADGNLWvaLWGGGGVVRFDPDGEL 218
|
|
| EGF_2 |
pfam07974 |
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins. |
250-272 |
6.20e-04 |
|
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
Pssm-ID: 400365 Cd Length: 26 Bit Score: 38.87 E-value: 6.20e-04
|
| YD_repeat_2x |
TIGR01643 |
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular ... |
1551-1591 |
7.71e-04 |
|
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.
Pssm-ID: 273728 [Multi-domain] Cd Length: 42 Bit Score: 39.11 E-value: 7.71e-04
10 20 30 40
....*....|....*....|....*....|....*....|..
gi 2421800221 1551 YSSTGQ-IASIQRGTTSEKVDYDGQGRIVSRVFADGKTWSYT 1591
Cdd:TIGR01643 1 YDAAGRlTGSTDADGTTTRYTYDAAGRLVEITDADGGSTRYE 42
|
|
| NHL_like_5 |
cd14963 |
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ... |
899-1149 |
9.62e-04 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271333 [Multi-domain] Cd Length: 268 Bit Score: 43.43 E-value: 9.62e-04
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 899 KLLAPVALACGIDGSLYVGDFnYVRRI--F-PSGNVTSVLelssnpAHRYYL--ATDPVT-----GDLYVSDTNTRRIYr 968
Cdd:cd14963 54 EFKYPYGIAVDSDGNIYVADL-YNGRIqvFdPDGKFLKYF------PEKKDRvkLISPAGlaiddGKLYVSDVKKHKVI- 125
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 969 pksltgakdltknaeVVAGTGEQCLPFdearcGDGGKAvEATLMSPKGMAVDKNGLIYFVD--GTMIRKVDQNG-IISTL 1045
Cdd:cd14963 126 ---------------VFDLEGKLLLEF-----GKPGSE-PGELSYPNGIAVDEDGNIYVADsgNGRIQVFDKNGkFIKEL 184
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1046 LGSNDLTSArpltcdtsmhisqvrLEWPTDLAINPmDNSIYVLDN--NVVLQITENRQVRIAAGRpmhcqvPGVEypvgk 1123
Cdd:cd14963 185 NGSPDGKSG---------------FVNPRGIAVDP-DGNLYVVDNlsHRVYVFDEQGKELFTFGG------RGKD----- 237
|
250 260
....*....|....*....|....*.
gi 2421800221 1124 havQTTLESATAIAVSYSGVLYITET 1149
Cdd:cd14963 238 ---DGQFNLPNGLFIDDDGRLYVTDR 260
|
|
| EGF_2 |
pfam07974 |
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins. |
380-404 |
1.36e-03 |
|
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
Pssm-ID: 400365 Cd Length: 26 Bit Score: 38.10 E-value: 1.36e-03
|
| EGF_2 |
pfam07974 |
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins. |
347-369 |
2.38e-03 |
|
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
Pssm-ID: 400365 Cd Length: 26 Bit Score: 37.33 E-value: 2.38e-03
|
| NHL_like_5 |
cd14963 |
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ... |
1198-1294 |
2.66e-03 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271333 [Multi-domain] Cd Length: 268 Bit Score: 42.28 E-value: 2.66e-03
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1198 KLSAPSSLAASPDGTLYIADLGNIRIRAVSKN----------KPLLNSM----------NFYeVASPTDQELYIFDINGT 1257
Cdd:cd14963 54 EFKYPYGIAVDSDGNIYVADLYNGRIQVFDPDgkflkyfpekKDRVKLIspaglaiddgKLY-VSDVKKHKVIVFDLEGK 132
|
90 100 110 120
....*....|....*....|....*....|....*....|...
gi 2421800221 1258 HQYTVSLVtGDYLYNFSYSN------DNDItAVTDSNGNtlRI 1294
Cdd:cd14963 133 LLLEFGKP-GSEPGELSYPNgiavdeDGNI-YVADSGNG--RI 171
|
|
| EGF_CA |
cd00054 |
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular ... |
472-501 |
3.88e-03 |
|
Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Pssm-ID: 238011 Cd Length: 38 Bit Score: 36.85 E-value: 3.88e-03
10 20 30
....*....|....*....|....*....|
gi 2421800221 472 PGLCNSNGRCTLDQNGWHCVCQPGWRGAGC 501
Cdd:cd00054 8 GNPCQNGGTCVNTVGSYRCSCPPGYTGRNC 37
|
|
| NHL_like_6 |
cd14962 |
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) ... |
1004-1223 |
5.15e-03 |
|
Uncharacterized NHL-repeat domain in bacterial proteins; The NHL (NCL-1, HT2A and LIN-41) repeat is found in multiple tandem copies, typically as 6 instances. It is about 40 residues long and resembles the WD repeat and other beta-propeller structures.
Pssm-ID: 271332 [Multi-domain] Cd Length: 271 Bit Score: 41.42 E-value: 5.15e-03
10 20 30 40 50 60 70 80
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1004 GKAVEATLMSPKGMAVDKNGLIYFVDGT--MIRKVDQNGIISTLLGSNDLtsarpltcdtsmhisQVRlewPTDLAINPM 1081
Cdd:cd14962 49 GNAGPNRFVSPIGVAIDANGNLYVSDAElgKVFVFDRDGKFLRAIGAGAL---------------FKR---PTGIAVDPA 110
|
90 100 110 120 130 140 150 160
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1082 DNSIYVLDnnvvlqiTENRQVRI--AAGRPMHcQVPgveyPVGKHAVQttLESATAIAVSYSGVLYITETDEKKINRI-- 1157
Cdd:cd14962 111 GKRLYVVD-------TLAHKVKVfdLDGRLLF-DIG----KRGSGPGE--FNLPTDLAVDRDGNLYVTDTMNFRVQIFda 176
|
170 180 190 200 210 220 230 240
....*....|....*....|....*....|....*....|....*....|....*....|....*....|....*....|
gi 2421800221 1158 --RQVTTDGEISLVAG---IPSECDCKNDAN---CDCYQS---------------GDGYAKDAKLSAPSSLAASPDGTLY 1214
Cdd:cd14962 177 dgKFLRSFGERGDGPGsfaRPKGIAVDSEGNiyvVDAAFDnvqifnpegellltvGGPGSGPGEFYLPSGIAIDKDDRIY 256
|
....*....
gi 2421800221 1215 IADLGNIRI 1223
Cdd:cd14962 257 VVDQFNRRI 265
|
|
| RHS_repeat |
pfam05593 |
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be ... |
1343-1374 |
7.47e-03 |
|
RHS Repeat; RHS proteins contain extended repeat regions. These repeats often appear to be involved in ligand binding. Note that this model may not find all the repeats in a protein and that it covers two RHS repeats. The 3D structure of an RHS-repeat-containing protein (the B and C components of an ABC toxin complex) has been determined. The RHS repeats form an extended strip of beta-sheet that spirals around to form a hollow shell, encapsulating the variable C-terminal domain.
Pssm-ID: 461685 [Multi-domain] Cd Length: 37 Bit Score: 36.04 E-value: 7.47e-03
10 20 30
....*....|....*....|....*....|..
gi 2421800221 1343 GLLATKSDETGWTTFFDYDSEGRLTNVTFPTG 1374
Cdd:pfam05593 5 GRLTSVTDPDGRVTTYTYDAAGRLTAVTDPDG 36
|
|
| YD_repeat_2x |
TIGR01643 |
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular ... |
1339-1377 |
9.31e-03 |
|
YD repeat (two copies); This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.
Pssm-ID: 273728 [Multi-domain] Cd Length: 42 Bit Score: 36.03 E-value: 9.31e-03
10 20 30
....*....|....*....|....*....|....*....
gi 2421800221 1339 HGNSGLLATKSDETGWTTFFDYDSEGRLTNVTFPTGVVT 1377
Cdd:TIGR01643 1 YDAAGRLTGSTDADGTTTRYTYDAAGRLVEITDADGGST 39
|
|
| EGF_2 |
pfam07974 |
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins. |
315-337 |
9.42e-03 |
|
EGF-like domain; This family contains EGF domains found in a variety of extracellular proteins.
Pssm-ID: 400365 Cd Length: 26 Bit Score: 35.79 E-value: 9.42e-03
|
| EGF |
pfam00008 |
EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very ... |
475-498 |
9.89e-03 |
|
EGF-like domain; There is no clear separation between noise and signal. pfam00053 is very similar, but has 8 instead of 6 conserved cysteines. Includes some cytokine receptors. The EGF domain misses the N-terminus regions of the Ca2+ binding EGF domains (this is the main reason of discrepancy between swiss-prot domain start/end and Pfam). The family is hard to model due to many similar but different sub-types of EGF domains. Pfam certainly misses a number of EGF domains.
Pssm-ID: 394967 Cd Length: 31 Bit Score: 35.82 E-value: 9.89e-03
|
|