Skip to content

Performance Issue with Extracting genres Attribute from Media #1449

Description

@sqkkyzx

Describe the Bug

For the following code snippet, I observed a significant performance discrepancy:

logging.info(F"movie count: {len(movies)}")

# test1
test1_start = time.time()
for movie in movies:
    logging.debug(f"{movie.__dict__.get('genres')}")
logging.info(f"duration {time.time()-test1_start}")

# test2
test2_start = time.time()
for movie in movies:
    logging.debug(f"{movie.genres}")
logging.info(f"duration {time.time()-test2_start}")

The output is:

INFO:root:movie count: 115
INFO:root:duration 0.0010001659393310547
INFO:root:duration 7.938525915145874

This is horrifying.

Code Snippets

No response

Expected Behavior

No response

Additional Context

No response

Operating System and Version

windows server 2022

Plex Media Server Version

1.40.4.8679

Python Version

3.12

PlexAPI Version

4.15.15

Activity

  1. JonnyWong16 commented on Aug 2, 2024

    @JonnyWong16
    Collaborator

    See responses in #713, #1007, #1128, and #1188 regarding partial objects.

    And autoreload configuration option.

    https://python-plexapi.readthedocs.io/en/latest/configuration.html#section-plexapi-options

  2. sqkkyzx commented on Aug 3, 2024

    @sqkkyzx
    Author

    Thank you for your reply. I completely understand that every time I call the media property, it requests all fields of that media from the Plex backend again. This is indeed a very good handling logic, and I agree with it.

    However, I still have some concerns about performance. The reason is that there is a significant difference between this approach and directly sending requests to PMS.

    Assuming I need to get the latest metadata for all media, I compared the following:

    • Test1: Iterate through the medias list and call the media.reload() method for each media.
    • Test2: Iterate through the medias list and construct the URL using f"{baseurl}{media.__dict__.get(key)}" for each media, then initiate a network request.
    • Test3: Run Test1 using 8 threads.
    • Test4: Run Test2 using 8 threads.

    The results are:

    INFO:root:msg="Media count: 802"
    INFO:root:msg="Test1" duration=84150ms
    INFO:root:msg="Test2" duration=5533ms
    INFO:root:msg="Test3" duration=33943ms
    INFO:root:msg="Test4" duration=1296ms
    

    My test code:

    import logging
    import time
    import concurrent.futures
    
    import requests
    from plexapi.myplex import PlexServer
    
    
    logging.basicConfig(level=logging.INFO)
    
    
    def threads(datalist, func, thread_count):
        def chunks(lst, n):
            for i in range(0, len(lst), n):
                yield lst[i:i + n]
    
        chunk_size = (len(datalist) + thread_count - 1) // thread_count
        list_chunks = list(chunks(datalist, chunk_size))
    
        with concurrent.futures.ThreadPoolExecutor(max_workers=thread_count) as executor:
            result_items = list(executor.map(func, [item for chunk in list_chunks for item in chunk]))
    
        return result_items
    
    
    def test(baseurl, token):
        client = PlexServer(baseurl, token)
    
        medias: list = client.library.all(libtype='movie')
    
        logging.info(f'msg="Media count: {len(medias)}"')
    
        t1 = int(time.time() * 1000)
        # test1
        for media in medias:
            metadata = media.reload().__dict__
            logging.debug(metadata)
        t2 = int(time.time() * 1000)
        logging.info(f'msg="Test1" duration={(t2 - t1)}ms')
    
        # test2
        for media in medias:
            metadata = requests.get(
                url=f'{baseurl}{media.__dict__.get('key')}', headers={'X-Plex-Token': token, 'Accept': 'application/json'}
            ).json().get("MediaContainer", {}).get("Metadata", [{}])[0]
            logging.debug(metadata)
        t3 = int(time.time() * 1000)
        logging.info(f'msg="Test2" duration={(t3 - t2)}ms')
    
        # test3
        def query_metadata_1(_media):
            _metadata = _media.reload().__dict__
            logging.debug(_metadata)
    
        threads(medias, query_metadata_1, 8)
        t4 = int(time.time() * 1000)
        logging.info(f'msg="Test3" duration={(t4 - t3)}ms')
    
        # test3
        def query_metadata_2(_media):
            _metadata = requests.get(
                url=f'{baseurl}{_media.__dict__.get('key')}', headers={'X-Plex-Token': token, 'Accept': 'application/json'}
            ).json().get("MediaContainer", {}).get("Metadata", [{}])[0]
            logging.debug(_metadata)
    
        threads(medias, query_metadata_2, 8)
        t5 = int(time.time() * 1000)
        logging.info(f'msg="Test4" duration={(t5 - t4)}ms')
    
    
    if __name__ == '__main__':
        test('http://127.0.0.1:32400', 'MyPlexToken')
    
    
  3. JonnyWong16 commented on Aug 3, 2024

    @JonnyWong16
    Collaborator

    Your direct query doesn't include all the URL parameters, so you are missing a lot of metadata.

    https://github.com/pkkid/python-plexapi/blob/d88f14e271828b9b435e0c8e819cf5b465eeb7a8/plexapi/base.py#L507-L527

    A more apples-to-apples comparison would be

            metadata = media.reload(
                checkFiles=False,
                includeAllConcerts=False,
                includeBandwidths=False,
                includeChapters=False,
                includeChildren=False,
                includeConcerts=False,
                includeExternalMedia=False,
                includeExtras=False,
                includeFields=False,
                includeGeolocation=False,
                includeLoudnessRamps=False,
                includeMarkers=False,
                includeOnDeck=False,
                includePopularLeaves=False,
                includePreferences=False,
                includeRelated=False,
                includeRelatedCount=False,
                includeReviews=False,
                includeStations=False
            ).__dict__

    vs.

            metadata = requests.get(
                url=f'{baseurl}{media.__dict__.get("key")}', headers={'X-Plex-Token': token, 'Accept': 'application/json'}
            ).json().get("MediaContainer", {}).get("Metadata", [{}])[0]

    Or

            metadata = media.reload().__dict__

    vs.

            metadata = requests.get(
                url=f'{baseurl}{media.__dict__.get("key")}',
                headers={'X-Plex-Token': token, 'Accept': 'application/json'},
                params=media._INCLUDES
            ).json().get("MediaContainer", {}).get("Metadata", [{}])[0]

    The remaining difference in performance will be the small overhead for building the PlexObject instead of just returning raw xml/json.

  4. sqkkyzx commented on Aug 3, 2024

    @sqkkyzx
    Author

    Understood. You are right; based on my tests, querying more metadata significantly increases the overhead. In a single-threaded environment, directly requesting the JSON is only 16ms faster on average per media item compared to constructing an object. However, with 8 threads, the difference is negligible.

    For my use case, I indeed do not need that much metadata. However, using

            media.reload(
                checkFiles=False,
                includeAllConcerts=False,
                includeBandwidths=False,
                includeChapters=False,
                includeChildren=False,
                includeConcerts=False,
                includeExternalMedia=False,
                includeExtras=False,
                includeFields=False,
                includeGeolocation=False,
                includeLoudnessRamps=False,
                includeMarkers=False,
                includeOnDeck=False,
                includePopularLeaves=False,
                includePreferences=False,
                includeRelated=False,
                includeRelatedCount=False,
                includeReviews=False,
                includeStations=False
            )
    

    seems overly cumbersome. But my problem is resolved. Thank you.

  5. JonnyWong16 commented on Aug 3, 2024

    @JonnyWong16
    Collaborator

    The biggest contributor is probably just checkFiles=False because then Plex won't scan your hard drive to check the file exists. The other include parameters should have very small impact on the speed.

    I think I will change the default to checkFiles=False and have the user include it explicitly if they need to check for the file.

  6. sqkkyzx commented on Aug 4, 2024

    @sqkkyzx
    Author

    I think the solution of changing the default values is quite good.

    (However, in my use case, I ultimately opted for the approach of concatenating the URL to send network requests and parse the JSON. Because even if there are only slight differences in a single request, when there are many items in the media library, the total time quickly expands from a few seconds to over a minute. Of course, perhaps it is completely unnecessary to be so obsessed with performance optimization. A friend of mine said that people really don't care how long a scheduled script takes to run. I think he is right; in fact, when I'm actually using it, I don't care about the runtime either. It's just that during the process of writing the script, I still feel troubled by it.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions