Repository navigation
Understanding Urlparser
There are two kinds of links on a site:
-
direct links to a stream (
.mp4,.m3u8,.mpd) - they go to the player as they are - hoster links (VOE, Filemoon, Dood, YouTube, ok.ru, ...) - a page with a player; the stream appears only after JavaScript ran, ads were shown, a button was clicked
The urlparser (IPTVPlayer/libs/urlparser.py) finds the stream behind a hoster link without a browser. Hosters try hard to prevent exactly that, so their pages are obfuscated and change often. The good news: a resolver written once works for every host that links to this hoster.
# in getLinksForVideo(): only offer links the urlparser knows
if self.up.checkHostSupport(url) == 1:
linksTab.append({'name': self.up.getHostName(url, True), 'url': strwithmeta(url, {'Referer': pageUrl}), 'need_resolve': 1})
# in getVideoLinks(): when the user picks such a link
return self.up.getVideoLinkExt(url)| method | returns |
|---|---|
checkHostSupport(url) |
1 known, 0 unknown, -1 known as not supported |
getHostName(url, nameOnly=False) |
voe.sx, with nameOnly=True just voe
|
getVideoLinkExt(url) |
list of {'name': ..., 'url': ...} - the streams, best quality first. Empty when the resolver failed; the reason is then set as "last error" (e.g. Hosting "xyz.com" unknown.) |
urlparser.py holds two classes. urlparser has the big dict hostMap: domain → resolver function. pageParser holds the resolver functions themselves (parserVOESX, parserOKRU, ...).
self.hostMap = {
"1fichier.com": self.pp.parser1FICHIERCOM,
...
"filemoon.sx": self.pp.parserBYSE,
...
"voe.sx": self.pp.parserVOESX,
...
}To find the resolver, getParser() takes the domain of the url without www. (https://www.voe.sx/e/abc → voe.sx) and looks it up; if it is not there, it tries once more without the first part (player.example.com → example.com). The debug log shows this:
_________________getHostName: [https://newhoster.com/e/6pvw6ticrzlg] -> [newhoster.com]
urlparser.getParser II try host[newhoster.com]->host2[com]
Helpers in libs/urlparserhelper.py and tools/iptvtypes.py that every resolver uses:
-
urlparser.decorateUrl(url, {'User-Agent': ..., 'Referer': ..., 'Origin': ...})- attaches the headers the player must send (many CDNs answer 403 without the right Referer/UA) and detects the protocol (m3u8,mpd, ...) -
getDirectM3U8Playlist(url, sortWithMaxBitrate=99999999)- reads an HLS master playlist and returns one entry per quality, best first -
getMPDLinksWithMeta(url)- the same for DASH -
strwithmeta(url, {...})- a string that carries ametadict;url.meta.get('Referer')reads it back - subtitles:
'external_sub_tracks': [{'title': '', 'url': ..., 'lang': 'en'}]in the meta of the stream url
Most "unknown hoster" errors today are not new hosters, but new domains of known ones - hosters change domains every few weeks. Open the link in the browser: if the player page looks like one the urlparser knows (same player, same url scheme /e/<id>), adding the domain is enough:
"newdomain-of-voe.com": self.pp.parserVOESX,Many hosters are built on the same software (JWPlayer pages, often with packed JavaScript). For those there is the generic parserJWPLAYER - try it before you write a new resolver.
When no existing resolver fits, add a function to pageParser and an entry to hostMap. Before you start, look at ResolveURL - the Kodi resolvers are maintained very actively and usually show the current trick of a hoster.
A typical player page:
<script>
jwplayer("vplayer").setup({
sources: [{file: "https://cdn.newhoster.com/hls/abc/master.m3u8", label: "720p"}],
tracks: [{file: "https://newhoster.com/srt/abc_eng.vtt", label: "English", kind: "captions"}],
image: "https://newhoster.com/thumb/abc.jpg"
});
</script>The resolver:
def parserNEWHOSTER(self, baseUrl):
printDBG("parserNEWHOSTER baseUrl[%s]" % baseUrl)
HTTP_HEADER = self.cm.getDefaultHeader(browser='chrome')
# the page of the host that embedded the player, set by the host with strwithmeta
HTTP_HEADER['Referer'] = strwithmeta(baseUrl).meta.get('Referer', baseUrl)
# /f/<id> is the download page, /e/<id> the player
url = baseUrl.replace('/f/', '/e/')
sts, data = self.cm.getPage(url, {'header': HTTP_HEADER})
if not sts:
return []
subTracks = []
for subUrl, label in re.findall(r'''file:\s*["']([^"']+\.(?:vtt|srt))["'],\s*label:\s*["']([^"']+)["']''', data):
subTracks.append({'title': label, 'url': subUrl, 'lang': label[:2].lower()})
streamUrl = self.cm.ph.getSearchGroups(data, r'''sources:\s*\[\{\s*file:\s*["']([^"']+)["']''')[0]
if not streamUrl:
return []
host = urlparser.getDomain(url, False) # https://newhoster.com/
streamUrl = urlparser.decorateUrl(streamUrl, {'User-Agent': HTTP_HEADER['User-Agent'], 'Referer': host,
'Origin': host[:-1], 'external_sub_tracks': subTracks})
if '.m3u8' in streamUrl:
return getDirectM3U8Playlist(streamUrl, sortWithMaxBitrate=99999999)
return [{'name': 'MP4', 'url': streamUrl}]and in hostMap, in alphabetical order:
"newhoster.com": self.pp.parserNEWHOSTER,Rules for resolvers:
- HTTP only through
self.cm, Python 2 and 3 compatible (see Host standard) - return
[]when nothing was found - never a half url - always
decorateUrl()the stream with the User-Agent of the request and the Referer/Origin the hoster expects; the player uses exactly these headers - best quality first; one entry per quality, with a readable
name - no new domain entries without testing a real link of that domain
Obfuscated pages (escaped strings, packed JavaScript, code that has to run) are covered in Another new server in urlparser: typical strategies.
Based on the original text by Maxbambi, thank you!
Using E2iPlayer
- Install
- Controls and navigation
- Settings explained
- Playback and subtitles
- Playlists and selection
- Download manager
- Web interface
- Torrents with TorrServer
Captchas
Problems?
For developers