Author: admin-to-del

  • ‘Seems they rely too much on glue’

    ‘Seems they rely too much on glue’


    A California Cybertruck owner took his electric truck to the 114-degree desert and noticed signs of it falling apart.

    His Cybertruck’s light bar, attached with glue, began to detach as he was hand-washing his vehicle.

    What’s happening?

    As Torque News reported, the Cybertruck owner, Duncan, noticed the vehicle’s light bar hanging halfway off while cleaning it after a desert road trip.

    He was disappointed to discover his $100,000 electric vehicle falling apart and possibly unable to withstand high temperatures. Duncan has only ever hand-washed his Cybertruck and never takes it through an automatic car wash.

    He shared a photo of the partially detached light bar in a forum post to the Cybertruck Owners Club.

    In the caption, Duncan wrote, “Seems they rely too much on glue.” He asked forum members if anyone had encountered a similar situation.

    Why are high-quality EV materials important?

    In response to the forum post, Cybertruck owners agreed that they have been dissatisfied with Tesla’s use of glue in the EV’s assembly.

    “I have wondered why Tesla didn’t design the lightbar to be glued like it is but also with simple snap-in tabs on the ends that make it snap into place on the windshield,” one forum member commented.

    High-quality materials are essential in building new EVs because they support the vehicles’ safety, longevity, and performance.

    Especially given our planet’s steadily rising temperatures, EVs must be able to withstand high heat and extreme weather conditions. Drivers need to feel confident that their EVs are a durable and practical transportation choice worth the investment.

    Do you think the government should ban gas-powered lawn tools?

    No way

    Definitely

    Only certain tools

    I don’t know

    Click your choice to see results and speak your mind.

    Owning an EV is among the best ways to save money on gas and maintenance while reducing your reliance on dirty energy and limiting your pollution output.

    However, Tesla Cybertrucks and other EVs must be safe to drive and reasonably repairable and maintainable when damage occurs. By gaining consumers’ trust, EV manufacturers can help further the widespread adoption of EVs and contribute to cleaner, greener roads.

    What’s being done to improve EV construction?

    Technology companies have been working to enhance how EVs are built by discovering innovative advancements that improve them.

    For example, AkzoNobel’s Resicoat is a powder coating technology that protects EV battery cells and motors.

    Oak Ridge National Laboratory researchers developed advanced battery materials to speed up charge times and extend EV battery life.

    Whether you choose to buy an EV from Tesla or another auto brand, it’s fascinating to learn a bit about how your vehicle was built. It’s also highly recommended to contact your EV company if you notice unexpected signs of wear and tear so it can promptly remedy the situation if a manufacturing issue is to blame.

    Join our free newsletter for weekly updates on the latest innovations improving our lives and shaping our future, and don’t miss this cool list of easy ways to help yourself while helping the planet.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=ArZ5WF

  • ‘Seems they rely too much on glue’

    ‘Seems they rely too much on glue’


    A California Cybertruck owner took his electric truck to the 114-degree desert and noticed signs of it falling apart.

    His Cybertruck’s light bar, attached with glue, began to detach as he was hand-washing his vehicle.

    What’s happening?

    As Torque News reported, the Cybertruck owner, Duncan, noticed the vehicle’s light bar hanging halfway off while cleaning it after a desert road trip.

    He was disappointed to discover his $100,000 electric vehicle falling apart and possibly unable to withstand high temperatures. Duncan has only ever hand-washed his Cybertruck and never takes it through an automatic car wash.

    He shared a photo of the partially detached light bar in a forum post to the Cybertruck Owners Club.

    In the caption, Duncan wrote, “Seems they rely too much on glue.” He asked forum members if anyone had encountered a similar situation.

    Why are high-quality EV materials important?

    In response to the forum post, Cybertruck owners agreed that they have been dissatisfied with Tesla’s use of glue in the EV’s assembly.

    “I have wondered why Tesla didn’t design the lightbar to be glued like it is but also with simple snap-in tabs on the ends that make it snap into place on the windshield,” one forum member commented.

    High-quality materials are essential in building new EVs because they support the vehicles’ safety, longevity, and performance.

    Especially given our planet’s steadily rising temperatures, EVs must be able to withstand high heat and extreme weather conditions. Drivers need to feel confident that their EVs are a durable and practical transportation choice worth the investment.

    Do you think the government should ban gas-powered lawn tools?

    No way

    Definitely

    Only certain tools

    I don’t know

    Click your choice to see results and speak your mind.

    Owning an EV is among the best ways to save money on gas and maintenance while reducing your reliance on dirty energy and limiting your pollution output.

    However, Tesla Cybertrucks and other EVs must be safe to drive and reasonably repairable and maintainable when damage occurs. By gaining consumers’ trust, EV manufacturers can help further the widespread adoption of EVs and contribute to cleaner, greener roads.

    What’s being done to improve EV construction?

    Technology companies have been working to enhance how EVs are built by discovering innovative advancements that improve them.

    For example, AkzoNobel’s Resicoat is a powder coating technology that protects EV battery cells and motors.

    Oak Ridge National Laboratory researchers developed advanced battery materials to speed up charge times and extend EV battery life.

    Whether you choose to buy an EV from Tesla or another auto brand, it’s fascinating to learn a bit about how your vehicle was built. It’s also highly recommended to contact your EV company if you notice unexpected signs of wear and tear so it can promptly remedy the situation if a manufacturing issue is to blame.

    Join our free newsletter for weekly updates on the latest innovations improving our lives and shaping our future, and don’t miss this cool list of easy ways to help yourself while helping the planet.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=GtYD1

  • ‘Seems they rely too much on glue’

    ‘Seems they rely too much on glue’


    A California Cybertruck owner took his electric truck to the 114-degree desert and noticed signs of it falling apart.

    His Cybertruck’s light bar, attached with glue, began to detach as he was hand-washing his vehicle.

    What’s happening?

    As Torque News reported, the Cybertruck owner, Duncan, noticed the vehicle’s light bar hanging halfway off while cleaning it after a desert road trip.

    He was disappointed to discover his $100,000 electric vehicle falling apart and possibly unable to withstand high temperatures. Duncan has only ever hand-washed his Cybertruck and never takes it through an automatic car wash.

    He shared a photo of the partially detached light bar in a forum post to the Cybertruck Owners Club.

    In the caption, Duncan wrote, “Seems they rely too much on glue.” He asked forum members if anyone had encountered a similar situation.

    Why are high-quality EV materials important?

    In response to the forum post, Cybertruck owners agreed that they have been dissatisfied with Tesla’s use of glue in the EV’s assembly.

    “I have wondered why Tesla didn’t design the lightbar to be glued like it is but also with simple snap-in tabs on the ends that make it snap into place on the windshield,” one forum member commented.

    High-quality materials are essential in building new EVs because they support the vehicles’ safety, longevity, and performance.

    Especially given our planet’s steadily rising temperatures, EVs must be able to withstand high heat and extreme weather conditions. Drivers need to feel confident that their EVs are a durable and practical transportation choice worth the investment.

    Do you think the government should ban gas-powered lawn tools?

    No way

    Definitely

    Only certain tools

    I don’t know

    Click your choice to see results and speak your mind.

    Owning an EV is among the best ways to save money on gas and maintenance while reducing your reliance on dirty energy and limiting your pollution output.

    However, Tesla Cybertrucks and other EVs must be safe to drive and reasonably repairable and maintainable when damage occurs. By gaining consumers’ trust, EV manufacturers can help further the widespread adoption of EVs and contribute to cleaner, greener roads.

    What’s being done to improve EV construction?

    Technology companies have been working to enhance how EVs are built by discovering innovative advancements that improve them.

    For example, AkzoNobel’s Resicoat is a powder coating technology that protects EV battery cells and motors.

    Oak Ridge National Laboratory researchers developed advanced battery materials to speed up charge times and extend EV battery life.

    Whether you choose to buy an EV from Tesla or another auto brand, it’s fascinating to learn a bit about how your vehicle was built. It’s also highly recommended to contact your EV company if you notice unexpected signs of wear and tear so it can promptly remedy the situation if a manufacturing issue is to blame.

    Join our free newsletter for weekly updates on the latest innovations improving our lives and shaping our future, and don’t miss this cool list of easy ways to help yourself while helping the planet.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=ca2UVyy

  • ‘Seems they rely too much on glue’

    ‘Seems they rely too much on glue’


    A California Cybertruck owner took his electric truck to the 114-degree desert and noticed signs of it falling apart.

    His Cybertruck’s light bar, attached with glue, began to detach as he was hand-washing his vehicle.

    What’s happening?

    As Torque News reported, the Cybertruck owner, Duncan, noticed the vehicle’s light bar hanging halfway off while cleaning it after a desert road trip.

    He was disappointed to discover his $100,000 electric vehicle falling apart and possibly unable to withstand high temperatures. Duncan has only ever hand-washed his Cybertruck and never takes it through an automatic car wash.

    He shared a photo of the partially detached light bar in a forum post to the Cybertruck Owners Club.

    In the caption, Duncan wrote, “Seems they rely too much on glue.” He asked forum members if anyone had encountered a similar situation.

    Why are high-quality EV materials important?

    In response to the forum post, Cybertruck owners agreed that they have been dissatisfied with Tesla’s use of glue in the EV’s assembly.

    “I have wondered why Tesla didn’t design the lightbar to be glued like it is but also with simple snap-in tabs on the ends that make it snap into place on the windshield,” one forum member commented.

    High-quality materials are essential in building new EVs because they support the vehicles’ safety, longevity, and performance.

    Especially given our planet’s steadily rising temperatures, EVs must be able to withstand high heat and extreme weather conditions. Drivers need to feel confident that their EVs are a durable and practical transportation choice worth the investment.

    Do you think the government should ban gas-powered lawn tools?

    No way

    Definitely

    Only certain tools

    I don’t know

    Click your choice to see results and speak your mind.

    Owning an EV is among the best ways to save money on gas and maintenance while reducing your reliance on dirty energy and limiting your pollution output.

    However, Tesla Cybertrucks and other EVs must be safe to drive and reasonably repairable and maintainable when damage occurs. By gaining consumers’ trust, EV manufacturers can help further the widespread adoption of EVs and contribute to cleaner, greener roads.

    What’s being done to improve EV construction?

    Technology companies have been working to enhance how EVs are built by discovering innovative advancements that improve them.

    For example, AkzoNobel’s Resicoat is a powder coating technology that protects EV battery cells and motors.

    Oak Ridge National Laboratory researchers developed advanced battery materials to speed up charge times and extend EV battery life.

    Whether you choose to buy an EV from Tesla or another auto brand, it’s fascinating to learn a bit about how your vehicle was built. It’s also highly recommended to contact your EV company if you notice unexpected signs of wear and tear so it can promptly remedy the situation if a manufacturing issue is to blame.

    Join our free newsletter for weekly updates on the latest innovations improving our lives and shaping our future, and don’t miss this cool list of easy ways to help yourself while helping the planet.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=KBNJ

  • Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!

    Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!


    The big horror news of the week was that Halloween: The Game, an official video game based on John Carpenter’s original horror classic, is headed our way next year from Illfonic, and it will feature online multiplayer, offline bots, and a single-player story mode!

    The news gets even more exciting as we head into the weekend, as the team announced that original Shape actor Nick Castle performed the motion capture for the game’s Michael Myers!

    The official Halloween Portal account tweets, “Although it may be slightly different than the one he wore on film in 1978, Nick Castle put back on a suit to deliver the iconic walk, movements, and kills as only the OG Michael Myers could.

    “In what is just one of the many ways Compass and Illfonic are striving to bring the utmost authenticity to this game, we wouldn’t dream of having anyone other than Nick bring this character to life in this new digital realm and we are thrilled he was eager to participate!”

    Check out a behind the scenes shot of Castle performing the mo-cap below!

    Halloween: The Game will invite players to step into the chilling world of John Carpenter’s genre-defining film, now transformed into a suspenseful one-versus-many stealth horror experience. Put on the iconic mask to become the ultimate slasher, Michael Myers, stalking and executing the citizens of Haddonfield one by one, or striving to thwart Michael Myers’ plans as Civilians determined to save the unaware townsfolk before it’s too late.

    Halloween: The Game will unleash Michael Myers onto Xbox Series X|S, PlayStation 5, and PC via Steam and the Epic Games Store in 2026. Here’s everything you need to know.



    Source link

  • Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!

    Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!


    The big horror news of the week was that Halloween: The Game, an official video game based on John Carpenter’s original horror classic, is headed our way next year from Illfonic, and it will feature online multiplayer, offline bots, and a single-player story mode!

    The news gets even more exciting as we head into the weekend, as the team announced that original Shape actor Nick Castle performed the motion capture for the game’s Michael Myers!

    The official Halloween Portal account tweets, “Although it may be slightly different than the one he wore on film in 1978, Nick Castle put back on a suit to deliver the iconic walk, movements, and kills as only the OG Michael Myers could.

    “In what is just one of the many ways Compass and Illfonic are striving to bring the utmost authenticity to this game, we wouldn’t dream of having anyone other than Nick bring this character to life in this new digital realm and we are thrilled he was eager to participate!”

    Check out a behind the scenes shot of Castle performing the mo-cap below!

    Halloween: The Game will invite players to step into the chilling world of John Carpenter’s genre-defining film, now transformed into a suspenseful one-versus-many stealth horror experience. Put on the iconic mask to become the ultimate slasher, Michael Myers, stalking and executing the citizens of Haddonfield one by one, or striving to thwart Michael Myers’ plans as Civilians determined to save the unaware townsfolk before it’s too late.

    Halloween: The Game will unleash Michael Myers onto Xbox Series X|S, PlayStation 5, and PC via Steam and the Epic Games Store in 2026. Here’s everything you need to know.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=uX57

  • Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!

    Original Shape Actor Nick Castle Is Back as Michael Myers in ‘Halloween: The Game’!


    The big horror news of the week was that Halloween: The Game, an official video game based on John Carpenter’s original horror classic, is headed our way next year from Illfonic, and it will feature online multiplayer, offline bots, and a single-player story mode!

    The news gets even more exciting as we head into the weekend, as the team announced that original Shape actor Nick Castle performed the motion capture for the game’s Michael Myers!

    The official Halloween Portal account tweets, “Although it may be slightly different than the one he wore on film in 1978, Nick Castle put back on a suit to deliver the iconic walk, movements, and kills as only the OG Michael Myers could.

    “In what is just one of the many ways Compass and Illfonic are striving to bring the utmost authenticity to this game, we wouldn’t dream of having anyone other than Nick bring this character to life in this new digital realm and we are thrilled he was eager to participate!”

    Check out a behind the scenes shot of Castle performing the mo-cap below!

    Halloween: The Game will invite players to step into the chilling world of John Carpenter’s genre-defining film, now transformed into a suspenseful one-versus-many stealth horror experience. Put on the iconic mask to become the ultimate slasher, Michael Myers, stalking and executing the citizens of Haddonfield one by one, or striving to thwart Michael Myers’ plans as Civilians determined to save the unaware townsfolk before it’s too late.

    Halloween: The Game will unleash Michael Myers onto Xbox Series X|S, PlayStation 5, and PC via Steam and the Epic Games Store in 2026. Here’s everything you need to know.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=Yt16

  • Apple trained an LLM to efficiently understand long-form video

    Apple trained an LLM to efficiently understand long-form video


    Apple researchers have developed an adapted version of the SlowFast-LLaVA model that beats larger models at long-form video analysis and understanding. Here’s what that means.

    The nerdy bits

    Very basically, when an LLM is trained to also understand video, it learns to split videos into frames, apply computer vision to extract visual features, analyze how those features change over time, and align all of that with language so it can describe or reason about the video in the form of text.

    One very inefficient way to do this is to analyze every single frame of a video, which creates an overwhelming amount of duplicated information, since most frames rarely include significant changes from one to the next.

    With this overwhelming amount of duplicated information at hand, it is very easy to blow past the LLM’s context window, which is the maximum amount of information it can retain at once. Once an LLM exceeds its context window, in order for a conversation to keep going, it stops taking older tokens into account to make room for new ones as it predicts each new token.

    Of course, there are more efficient ways to train video LLMs (NVIDIA recently published an interesting paper on this), but this is the general idea to keep in mind for Apple’s study.

    Apple’s study

    As Apple’s researchers explain it in the paper SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding:

    “Video large language models (LLMs) integrate video perception into pre-trained LLMs to process videos and generate responses to user commands. Although significant progress has been made, notable limitations remain in existing Video LLMs.”

    The limitations, according to them, are threefold:

    • Existing models tend to rely heavily on long context windows and huge numbers of frames, which is inefficient and not easily transferable to smaller models;
    • Most of them require complex multi-stage training pipelines (often using private datasets) that are hard to reproduce;
    • Many are optimized only for video tasks, which limits their usefulness as general-purpose models that also understand images.

    To address those limitations, Apple first looked at SlowFast-LLaVA, an open-source model that had already shown promising results by combining spatial and temporal cues through a two-stream setup: a slow stream that looks at fewer frames in higher detail to capture what’s in the scene, and a fast stream that looks at more frames in lower detail to track how things move over time.

    First, Apple fine-tuned SlowFast-LLaVA on images, in order to build general visual reasoning capabilities. Then, it was trained jointly on both images and videos (from public datasets), to learn temporal structure without sacrificing image understanding.

    Image: Apple

    The result was SlowFast-LLaVA-1.5 (or SF-LLaVA-1.5), a family of models at 1B, 3B, and 7B parameter scales, that manages to outperform much larger models across a range of video tasks, sometimes “by significant margins,” as noted by the researchers themselves.

    Image: Apple

    In fact, on long-form video benchmarks like LongVideoBench and MLVU, Apple’s model sets new state-of-the-art results across all model sizes, including its smallest, 1B, version.

    What’s more, the model also overcomes one of the three shortcomings noted by the researchers, and performs well on image tasks too, including benchmarks for knowledge, math reasoning, OCR, and text-rich scenarios.

    Image: Apple

    The team even tested several video compression strategies, but found that their setup struck the best balance between speed, accuracy, and token count.

    Still, there are limitations

    With SF-LLaVA-1.5, Apple’s researchers decided that the model would have a maximum input frame length of 128.

    This means that whether it is analyzing a clip that is a few minutes or a few hours long, it always maxes out at 128 frames, with 96 evenly spaced frames selected for the fast stream, and 32 evenly spaced frames selected for the slow stream.

    With that in mind, the researchers say that:

    “This approach may miss some key frames in long-form videos and mislead the model about a video’s playback speed. (…) SF-LLaVA-1.5’s performance can be further improved by tuning all parameters, including the visual encoder. However, we found this is not trivial for Long Video LLMs due to the high GPU memory cost of caching the activation values. Future studies could explore the integration of memory-saving techniques, such as Stochastic BP.”

    That said, Apple’s approach rendered it a state-of-the-art model, with the extra chops of being trained exclusively on public datasets. SF-LLaVA-1.5 is now an open-source model available on GitHub and Hugging Face, and you can find the complete study on arXiv.

    Below are a few examples of the model in action:

    Image: Apple
    Image: Apple
    Image: Apple

    Limited time Apple Watch deals on Amazon

    FTC: We use income earning auto affiliate links. More.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=7VYYl

  • Apple trained an LLM to efficiently understand long-form video

    Apple trained an LLM to efficiently understand long-form video


    Apple researchers have developed an adapted version of the SlowFast-LLaVA model that beats larger models at long-form video analysis and understanding. Here’s what that means.

    The nerdy bits

    Very basically, when an LLM is trained to also understand video, it learns to split videos into frames, apply computer vision to extract visual features, analyze how those features change over time, and align all of that with language so it can describe or reason about the video in the form of text.

    One very inefficient way to do this is to analyze every single frame of a video, which creates an overwhelming amount of duplicated information, since most frames rarely include significant changes from one to the next.

    With this overwhelming amount of duplicated information at hand, it is very easy to blow past the LLM’s context window, which is the maximum amount of information it can retain at once. Once an LLM exceeds its context window, in order for a conversation to keep going, it stops taking older tokens into account to make room for new ones as it predicts each new token.

    Of course, there are more efficient ways to train video LLMs (NVIDIA recently published an interesting paper on this), but this is the general idea to keep in mind for Apple’s study.

    Apple’s study

    As Apple’s researchers explain it in the paper SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding:

    “Video large language models (LLMs) integrate video perception into pre-trained LLMs to process videos and generate responses to user commands. Although significant progress has been made, notable limitations remain in existing Video LLMs.”

    The limitations, according to them, are threefold:

    • Existing models tend to rely heavily on long context windows and huge numbers of frames, which is inefficient and not easily transferable to smaller models;
    • Most of them require complex multi-stage training pipelines (often using private datasets) that are hard to reproduce;
    • Many are optimized only for video tasks, which limits their usefulness as general-purpose models that also understand images.

    To address those limitations, Apple first looked at SlowFast-LLaVA, an open-source model that had already shown promising results by combining spatial and temporal cues through a two-stream setup: a slow stream that looks at fewer frames in higher detail to capture what’s in the scene, and a fast stream that looks at more frames in lower detail to track how things move over time.

    First, Apple fine-tuned SlowFast-LLaVA on images, in order to build general visual reasoning capabilities. Then, it was trained jointly on both images and videos (from public datasets), to learn temporal structure without sacrificing image understanding.

    Image: Apple

    The result was SlowFast-LLaVA-1.5 (or SF-LLaVA-1.5), a family of models at 1B, 3B, and 7B parameter scales, that manages to outperform much larger models across a range of video tasks, sometimes “by significant margins,” as noted by the researchers themselves.

    Image: Apple

    In fact, on long-form video benchmarks like LongVideoBench and MLVU, Apple’s model sets new state-of-the-art results across all model sizes, including its smallest, 1B, version.

    What’s more, the model also overcomes one of the three shortcomings noted by the researchers, and performs well on image tasks too, including benchmarks for knowledge, math reasoning, OCR, and text-rich scenarios.

    Image: Apple

    The team even tested several video compression strategies, but found that their setup struck the best balance between speed, accuracy, and token count.

    Still, there are limitations

    With SF-LLaVA-1.5, Apple’s researchers decided that the model would have a maximum input frame length of 128.

    This means that whether it is analyzing a clip that is a few minutes or a few hours long, it always maxes out at 128 frames, with 96 evenly spaced frames selected for the fast stream, and 32 evenly spaced frames selected for the slow stream.

    With that in mind, the researchers say that:

    “This approach may miss some key frames in long-form videos and mislead the model about a video’s playback speed. (…) SF-LLaVA-1.5’s performance can be further improved by tuning all parameters, including the visual encoder. However, we found this is not trivial for Long Video LLMs due to the high GPU memory cost of caching the activation values. Future studies could explore the integration of memory-saving techniques, such as Stochastic BP.”

    That said, Apple’s approach rendered it a state-of-the-art model, with the extra chops of being trained exclusively on public datasets. SF-LLaVA-1.5 is now an open-source model available on GitHub and Hugging Face, and you can find the complete study on arXiv.

    Below are a few examples of the model in action:

    Image: Apple
    Image: Apple
    Image: Apple

    Limited time Apple Watch deals on Amazon

    FTC: We use income earning auto affiliate links. More.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=LVqt

  • Apple trained an LLM to efficiently understand long-form video

    Apple trained an LLM to efficiently understand long-form video


    Apple researchers have developed an adapted version of the SlowFast-LLaVA model that beats larger models at long-form video analysis and understanding. Here’s what that means.

    The nerdy bits

    Very basically, when an LLM is trained to also understand video, it learns to split videos into frames, apply computer vision to extract visual features, analyze how those features change over time, and align all of that with language so it can describe or reason about the video in the form of text.

    One very inefficient way to do this is to analyze every single frame of a video, which creates an overwhelming amount of duplicated information, since most frames rarely include significant changes from one to the next.

    With this overwhelming amount of duplicated information at hand, it is very easy to blow past the LLM’s context window, which is the maximum amount of information it can retain at once. Once an LLM exceeds its context window, in order for a conversation to keep going, it stops taking older tokens into account to make room for new ones as it predicts each new token.

    Of course, there are more efficient ways to train video LLMs (NVIDIA recently published an interesting paper on this), but this is the general idea to keep in mind for Apple’s study.

    Apple’s study

    As Apple’s researchers explain it in the paper SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding:

    “Video large language models (LLMs) integrate video perception into pre-trained LLMs to process videos and generate responses to user commands. Although significant progress has been made, notable limitations remain in existing Video LLMs.”

    The limitations, according to them, are threefold:

    • Existing models tend to rely heavily on long context windows and huge numbers of frames, which is inefficient and not easily transferable to smaller models;
    • Most of them require complex multi-stage training pipelines (often using private datasets) that are hard to reproduce;
    • Many are optimized only for video tasks, which limits their usefulness as general-purpose models that also understand images.

    To address those limitations, Apple first looked at SlowFast-LLaVA, an open-source model that had already shown promising results by combining spatial and temporal cues through a two-stream setup: a slow stream that looks at fewer frames in higher detail to capture what’s in the scene, and a fast stream that looks at more frames in lower detail to track how things move over time.

    First, Apple fine-tuned SlowFast-LLaVA on images, in order to build general visual reasoning capabilities. Then, it was trained jointly on both images and videos (from public datasets), to learn temporal structure without sacrificing image understanding.

    Image: Apple

    The result was SlowFast-LLaVA-1.5 (or SF-LLaVA-1.5), a family of models at 1B, 3B, and 7B parameter scales, that manages to outperform much larger models across a range of video tasks, sometimes “by significant margins,” as noted by the researchers themselves.

    Image: Apple

    In fact, on long-form video benchmarks like LongVideoBench and MLVU, Apple’s model sets new state-of-the-art results across all model sizes, including its smallest, 1B, version.

    What’s more, the model also overcomes one of the three shortcomings noted by the researchers, and performs well on image tasks too, including benchmarks for knowledge, math reasoning, OCR, and text-rich scenarios.

    Image: Apple

    The team even tested several video compression strategies, but found that their setup struck the best balance between speed, accuracy, and token count.

    Still, there are limitations

    With SF-LLaVA-1.5, Apple’s researchers decided that the model would have a maximum input frame length of 128.

    This means that whether it is analyzing a clip that is a few minutes or a few hours long, it always maxes out at 128 frames, with 96 evenly spaced frames selected for the fast stream, and 32 evenly spaced frames selected for the slow stream.

    With that in mind, the researchers say that:

    “This approach may miss some key frames in long-form videos and mislead the model about a video’s playback speed. (…) SF-LLaVA-1.5’s performance can be further improved by tuning all parameters, including the visual encoder. However, we found this is not trivial for Long Video LLMs due to the high GPU memory cost of caching the activation values. Future studies could explore the integration of memory-saving techniques, such as Stochastic BP.”

    That said, Apple’s approach rendered it a state-of-the-art model, with the extra chops of being trained exclusively on public datasets. SF-LLaVA-1.5 is now an open-source model available on GitHub and Hugging Face, and you can find the complete study on arXiv.

    Below are a few examples of the model in action:

    Image: Apple
    Image: Apple
    Image: Apple

    Limited time Apple Watch deals on Amazon

    FTC: We use income earning auto affiliate links. More.



    Source link

    https://d3n8a8pro7vhmx.cloudfront.net/alize/pages/34/attachments/original/1748981808/w.xml?o2x=GdKawvM