Skip to content

Understanding Urlparser

Masta2002 edited this page Oct 4, 2026 · 2 revisions

What the urlparser does

There are two kinds of links on a site:

  • direct links to a stream (.mp4, .m3u8, .mpd) - they go to the player as they are
  • hoster links (VOE, Filemoon, Dood, YouTube, ok.ru, ...) - a page with a player; the stream appears only after JavaScript ran, ads were shown, a button was clicked

The urlparser (IPTVPlayer/libs/urlparser.py) finds the stream behind a hoster link without a browser. Hosters try hard to prevent exactly that, so their pages are obfuscated and change often. The good news: a resolver written once works for every host that links to this hoster.

How a host uses it

# in getLinksForVideo(): only offer links the urlparser knows
if self.up.checkHostSupport(url) == 1:
    linksTab.append({'name': self.up.getHostName(url, True), 'url': strwithmeta(url, {'Referer': pageUrl}), 'need_resolve': 1})

# in getVideoLinks(): when the user picks such a link
return self.up.getVideoLinkExt(url)
method returns
checkHostSupport(url) 1 known, 0 unknown, -1 known as not supported
getHostName(url, nameOnly=False) voe.sx, with nameOnly=True just voe
getVideoLinkExt(url) list of {'name': ..., 'url': ...} - the streams, best quality first. Empty when the resolver failed; the reason is then set as "last error" (e.g. Hosting "xyz.com" unknown.)

How it is built

urlparser.py holds two classes. urlparser has the big dict hostMap: domain → resolver function. pageParser holds the resolver functions themselves (parserVOESX, parserOKRU, ...).

self.hostMap = {
    "1fichier.com": self.pp.parser1FICHIERCOM,
    ...
    "filemoon.sx": self.pp.parserBYSE,
    ...
    "voe.sx": self.pp.parserVOESX,
    ...
}

To find the resolver, getParser() takes the domain of the url without www. (https://www.voe.sx/e/abc → voe.sx) and looks it up; if it is not there, it tries once more without the first part (player.example.com → example.com). The debug log shows this:

_________________getHostName: [https://newhoster.com/e/6pvw6ticrzlg] -> [newhoster.com]
urlparser.getParser II try host[newhoster.com]->host2[com]

Helpers in libs/urlparserhelper.py and tools/iptvtypes.py that every resolver uses:

  • urlparser.decorateUrl(url, {'User-Agent': ..., 'Referer': ..., 'Origin': ...}) - attaches the headers the player must send (many CDNs answer 403 without the right Referer/UA) and detects the protocol (m3u8, mpd, ...)
  • getDirectM3U8Playlist(url, sortWithMaxBitrate=99999999) - reads an HLS master playlist and returns one entry per quality, best first
  • getMPDLinksWithMeta(url) - the same for DASH
  • strwithmeta(url, {...}) - a string that carries a meta dict; url.meta.get('Referer') reads it back
  • subtitles: 'external_sub_tracks': [{'title': '', 'url': ..., 'lang': 'en'}] in the meta of the stream url

A new domain of a known hoster

Most "unknown hoster" errors today are not new hosters, but new domains of known ones - hosters change domains every few weeks. Open the link in the browser: if the player page looks like one the urlparser knows (same player, same url scheme /e/<id>), adding the domain is enough:

            "newdomain-of-voe.com": self.pp.parserVOESX,

Many hosters are built on the same software (JWPlayer pages, often with packed JavaScript). For those there is the generic parserJWPLAYER - try it before you write a new resolver.

A new resolver

When no existing resolver fits, add a function to pageParser and an entry to hostMap. Before you start, look at ResolveURL - the Kodi resolvers are maintained very actively and usually show the current trick of a hoster.

A typical player page:

<script>
jwplayer("vplayer").setup({
    sources: [{file: "https://cdn.newhoster.com/hls/abc/master.m3u8", label: "720p"}],
    tracks: [{file: "https://newhoster.com/srt/abc_eng.vtt", label: "English", kind: "captions"}],
    image: "https://newhoster.com/thumb/abc.jpg"
});
</script>

The resolver:

    def parserNEWHOSTER(self, baseUrl):
        printDBG("parserNEWHOSTER baseUrl[%s]" % baseUrl)
        HTTP_HEADER = self.cm.getDefaultHeader(browser='chrome')
        # the page of the host that embedded the player, set by the host with strwithmeta
        HTTP_HEADER['Referer'] = strwithmeta(baseUrl).meta.get('Referer', baseUrl)
        # /f/<id> is the download page, /e/<id> the player
        url = baseUrl.replace('/f/', '/e/')
        sts, data = self.cm.getPage(url, {'header': HTTP_HEADER})
        if not sts:
            return []

        subTracks = []
        for subUrl, label in re.findall(r'''file:\s*["']([^"']+\.(?:vtt|srt))["'],\s*label:\s*["']([^"']+)["']''', data):
            subTracks.append({'title': label, 'url': subUrl, 'lang': label[:2].lower()})

        streamUrl = self.cm.ph.getSearchGroups(data, r'''sources:\s*\[\{\s*file:\s*["']([^"']+)["']''')[0]
        if not streamUrl:
            return []
        host = urlparser.getDomain(url, False)  # https://newhoster.com/
        streamUrl = urlparser.decorateUrl(streamUrl, {'User-Agent': HTTP_HEADER['User-Agent'], 'Referer': host,
                                                      'Origin': host[:-1], 'external_sub_tracks': subTracks})
        if '.m3u8' in streamUrl:
            return getDirectM3U8Playlist(streamUrl, sortWithMaxBitrate=99999999)
        return [{'name': 'MP4', 'url': streamUrl}]

and in hostMap, in alphabetical order:

            "newhoster.com": self.pp.parserNEWHOSTER,

Rules for resolvers:

  • HTTP only through self.cm, Python 2 and 3 compatible (see Host standard)
  • return [] when nothing was found - never a half url
  • always decorateUrl() the stream with the User-Agent of the request and the Referer/Origin the hoster expects; the player uses exactly these headers
  • best quality first; one entry per quality, with a readable name
  • no new domain entries without testing a real link of that domain

Obfuscated pages (escaped strings, packed JavaScript, code that has to run) are covered in Another new server in urlparser: typical strategies.


Based on the original text by Maxbambi, thank you!

Clone this wiki locally