Youtube dataset, contains duplicate video id's that are stored multiple times based on different keywords used. example: if i post a reaction video and tag it with music and reaction i get it twice in the dataset. this brigs issues regarding likes and views as it doubles them. What to do in this cenario ?
#๐ Youtube Dataset
8 messages ยท Page 1 of 1 (latest)
@junior harness
Remember to:
- Ask your Python question, not if you can ask or if there's an expert who can help.
- Show a code sample as text (rather than a screenshot) and the error message, if you've got one.
- Explain what you expect to happen and what actually happens.
:warning: Do not pip install anything that isn't related to your question, especially if asked to over DMs.
Closes after a period of inactivity, or when you send !close.
drop duplicates based on the ID subset and/or groupby and transform the Keyword column into a list instead of individual values
this dataset looks a lot like it was in a sane format (a few tables, including a pivot table mapping ids to keywords), and then someone made an outer join on it
This help channel has been closed and it's no longer possible to send messages here. If your question wasn't answered, feel free to create a new post in #1035199133436354600. To maximize your chances of getting a response, check out this guide on asking good questions.