Leveraging Multimodal LLM for Smarter Decision-Making in Specialized Fields
Artificial intelligence models can process different types of data, such as text, images, audio, and structured data. When a model combines multiple types of data, it is called a multimodal model. Unlike unimodal models that rely on just one data source, multimodal models merge different types of data to improve accuracy and produce better results.
These models are useful in real-world applications such as medical imaging, self-driving cars, and language understanding. By analyzing multiple forms of input, they gain a better understanding of the context, making their predictions more reliable and precise.

General AI models are trained in a wide range of data, but they may not perform well in specialized fields. When a model is fine-tuned using data specific to a particular industry, such as construction or healthcare, it learns patterns unique to that field. This improves its performance and ensures better accuracy in real-world applications.
For example, a model trained in general text might struggle to interpret complex engineering drawings. However, if the model is trained on domain-specific data like construction plans and cost estimation patterns, it can make more precise predictions and assist professionals in making informed decisions. Additionally, we can also train a quantized version of the model to deploy in our system, which helps reduce computational costs and avoid privacy concerns by keeping sensitive data within the deployment environment.
● Faster Processing: A streamlined model can generate results quickly, improving responsiveness.
● Enhanced Privacy and Security: Deploying a model locally ensures sensitive data remains secure, avoiding cloud-related privacy risks.
● Scalability: Lightweight models can run on various devices, making AI accessible for different applications, including mobile and edge computing.
LLaMA 3.2-11B is one of the most advanced multimodal AI models designed for understanding and generating human-like text while integrating additional data formats. With 11 billion parameters, it can provide detailed and accurate responses for various applications.
Key features of LLaMA 3.2-11B include:
● Fine-Tuning Flexibility: It can be adapted for specific industry tasks.
● Improved Accuracy: It learns patterns from large-scale data, leading to better results.
By customizing LLaMA 3.2-11B with domain-specific training, industries like construction, healthcare, and finance can enhance their decision-making processes.
To showcase the benefits of multimodal models, let’s examine how AI can improve the analysis of floor plan documents and property assessments. Traditionally, this process requires manual inspection, which is time-consuming and prone to errors. By using multimodal AI, we can automate document analysis by integrating text-based project descriptions with architectural plans and property data.
To showcase the benefits of multimodal models, let’s examine how AI can improve the analysis of floor plan documents and property assessments. Traditionally, this process requires manual inspection, which is time-consuming and prone to errors. By using multimodal AI, we can automate document analysis by integrating text-based project descriptions with architectural plans and property data.
● Dataset Source: Kaggle dataset with annotated images
o https://www.kaggle.com/datasets/adilmohammed/floor-plan-images-and-their-details
● Annotations: 1,000 images labeled with Japanese text description.
● Fine-Tuning Process: The model is trained to recognize and interpret Japanese text accurately.
● Input Prompt Format: English question with an image input, ensuring cross-language understanding.

When using a basic text-based AI model to analyze floor plans and properties without training it on industry-specific data, the following issues arise:
Lack of Domain-Specific Knowledge: The model was not trained on architectural or real estate datasets, leading to misinterpretations.
Inability to Handle Floor Plan Context: The model struggles to process and summarize structured information accurately.
Random Numerical Generation: Without fine-tuning, the model does not recognize numerical patterns relevant to measurements.
Textual Hallucinations: The model attempts to generate a "story" instead of providing factual, structured data.

The model used for this case study is Unsloth/Llama-3.2-11B-Vision-BnB-4bit.
Sample code used for this case study: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.2_(11B)-Vision.ipynb
この家のフロア・プランを見てみると、フロア計測は29m X 37m です。これはあまり一般的ではない家形ですが、バルコニーを持った広いアパートのような感じでしょうか。また、ここでは床が灰色で壁は白ですが、これもあまり一般的ではありませんが、現実の物に似せてみた形になります。
1階には「SHOP」と「LIVING RM」を、2階には「BEDROOM」があるようで、この2階の「BEDROOM」の大きさは 13m X 10m、13m^2で、12m X 10mのバルコニーに 2.0mの壁で囲まれたルームとなっており、これは通常の部屋の大きさよりもかなり大きな大きさです。つまり、ここは部屋と庭で構成される大きく分けられた構成物で、通常の部屋よりも多く取り入ることが可能です。そのため、ここにアッセムブリルなど大きな家具を置けるようにしたり、大きなトイレや浴槽の入るところにも気を付けたりすることが可能です。
1階には「SHOP」と「LIVING RM」があり、 SHOPには「29m X 37m」で、1階の大きさと同じ大きさです。
また、LIVING RMは17m X 12mで9.0m^2です。SHOPよりも1階より少し大きく、9.0m^2よりも小さい大きさのLIVING RMが入っています。これが全体的にフロア・プランに好感がもたれた理由の1つですね。
このフロア・プランは「SHOP」に「LIVING RM」のように、「BEDROOM」に「BATH」などのルームがあるように、現実の物に合わせたルームの入る場所と、それぞれの部屋のサイズは「SHOP」は 29m X 37m 、「LIVING RM」は17m X 12m「BEDROOM」は13m X 10m、10m^2「BATH」は7.0m X 2.0mです。LIVING RMの部分もここは少し広くなっていて、SHOPよりも少し広くなり、ここもまた良くある家に似ており、このフロア・プランはLIVING RMに一番似ていますね。この2階に「SHOP」の部分はかなりの広く、13m^2とあるからも、これがとてもよい特徴だと考えて良いです。LIVING RMも17m X 12mですが、この2階の「SHOP」部分はかなりの大きさであり、2階に「LIVING RM」の部分が「SHOP」と同じ大きさであってよいのです。
また、LIVING RM では「13m^2」とありますが、これがフロアプランの2階「SHOP」でとても大きな特徴があるからです。「BATH」と「BAR」は2階で10m^2で、7.0m^2のBATHと、12m^2のバルの合計12m^2の小さな大きさを持ち、バルコニーも大きくして2階にあり、これはここでもLIVING RMとSHOPの大きさを合わせると全てのフロア計測計が大きな特徴になっていることです。バルの大きさは小さくなっていますが、これはそこをフロア計測から無視することもあり得る、LIVING RMが12m^2でバルコニーもあり、2階は2mに囲まれるのです。
この2階「SHOP」部分は「29m X 37m」と「LIVING RM」と同じ計測計があり、SHOPは大きな特徴になっていることです。「BATH」は2階にあり、2階のフロア計測に一番大きい計測計で、7.0m X 2.0mの計測計を持っています。全体では、「BEDROOM」は2階にあり、2階では大きな大きさがあり、10m^2で、2階ではこれの大小比は一番大きい計測計です。2階の「LIVING RM」の大小比は、2階の「SHOP」よりも大きく、9.0m^2に、9.0m^2に同じ計測計と同じです。この2階の計測計の大小は「BATH」の計測計よりもかなり大きい、SHOPがとても良い大きさであることもあって、LIVING RMも全体で大きな計測計になったこともとても良いのです。
このフロアプランはフロア計測が29m X 37m です。これはフロア計測のサイズがかなり広いので、フロア計測をよく見ると、フロア計測には2階と1階がフロア計測してありますが、ここにアッセムブリルとBATHが入っています。
また、このフロア計測に2階にも1階にもある計測計は2階に2.0mの壁があります。2階にも2.0mの壁があるため、フロア計測を2階で2.0mの壁を考えてみると、ここはかなり広くなることがわかりますが、フロア計測の大小と2階の大きさが違いがあると考えます。また、1階と2階の計測計が同じでも、2階と1階が異なっているので、2階でとても大きい計測計でも、1階はあまり大きな計測計でもないことを考えてみて、1階があまり大きくないことはわかります。
フロア計測の大小は2階の大きさよりもはるか大きいですが、これをとても大きな計測計で考える必要があります。2階に1階はあまり広くないですが、2階でとても広い計測計であるからです。これが2階に大きな計測計として使い、フロア計測から見える計測計でもあるのです。したがって、これはフロア計測から2階の大きさと、1階の計測計を見るととても広い計測計になるということを考えてみると、このフロア計測は2階の計測計と1階の計測計が違いがあり、2階の計測計はとても大きな計測計でも、1階の計測計はあまり広くない計測計で2階に大きな計測計はあまり広くないことからです。
2階でLIVING RMと同じ計測計があり、これも全ての計測計に合わせて良い計測計にできると、ここは大きな計測計に大きな計測計にもなっているので、計測計がLIVING RMの計測計の大きさと同じ計測計です。
2階では計測計の大きさもLIVING RMの計測計の大きさと同じ計測計でも、全体でとても良い計測計でできてしまっていることから、この2階のフロア計測はフロア計測の計測計の大小に大きな計測計でも、フロア計測の計測計に全体で大きな計測計でも、LIVING RMの計測計もフロア計測の計測計の大小
に全体で大きな計測計で計測計の計測計の大小にも合わせて計測計と合わせても、計測計と合わせても全体で大きな計測計の計測計を合わせて計測計の計測計に合わせて計測計と合わせて計測計の計測計に合わせて計測計を合わせて計測計を合わせて計測計の計測計の計測計に合わせて計測計の
計測計の計測計を合わせて計測計と合わせて計測計を合わせて計測計の計測計を合わせて計測計の計測計の計測計を合わせて計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測計の計測計の計測計を合わせて計測計を合わせて計測計を合わせて計測計の計測計に合わせて計測計を
合わせて計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測
計を合わせて計測計を合わせて計測計の計測計の計測計に合わせて計測計を合わせて計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測計の計測計の計測計に合わせて計測計の計測計の計測
After fine-tuning the model on floor plan data, the results showed significant improvements:
● Accurate Room Detection and Labeling: The model correctly identified room types such as bedrooms, bathrooms, living areas, and shop spaces.
● Improved Measurement Precision: Room sizes were now correctly extracted and labeled, aligning closely with ground truth data.
● Enhanced Structural and Layout Understanding: The model recognized key features, including stair placements, windows, and functional areas.
● Proper Functional Area Allocation: The relationship between different areas was well-defined, such as the proximity of the living room to the bar and the connectivity between shop spaces and porches.
*部屋数:**
メインフロアには、1つのベッドルーム、1つのバスルーム、1つのリビングルーム、1つのショップスペースがあります。
**部屋の種類とラベル:**
ベッドルーム、バスルーム、リビングルーム、ショップスペースがあります。
**部屋のサイズ:**
- ベッドルーム:13' x 10'
- リビングルーム:17' x 12'
- ショップスペース:29' x 37'
**主な特徴:**
玄関の近くに階段、リビングルームに窓、ベッドルームに暖炉、バスルームに浴槽が含まれています。
**機能エリア:**
リビングルームの近くにバーがあります。ショップスペースにはポーチが接続されます。
**レイアウト概要:**
リビングスペースを通る中央の廊下。ベッドルームは右側にあり、バスルームはベッドルームに隣接しています。ショップスペースはメインフロアを横断しています。
**床材と特性:**
ショップスペースとリビングルームは明確な床材が見えますが、詳細は示されていません。
**概要:**
メインフロアにはショップスペース、ベッドルーム、リビングエリアを含み、オープンな流れでユーティリティエリアと居住エリアが統合されています。主要な特徴としては、バスルームに浴槽と階段の近さを含みます。
Demo Link:
Using multimodal AI models like LLaMA 3.2-11B enhances industry-specific applications such as analyzing floor plan documents and property assessments. By combining text, images, and structured data, these models provide deeper insights, leading to better accuracy and efficiency. These details can further be used to calculate cost estimation in the future based on client requirements to help select the best floor plan for a new house in less time. Additionally, in the future, material and labor costs can be calculated from these floor plan images, making the decision-making process more efficient and cost-effective. As AI advances, multimodal learning will play a crucial role in improving decision-making across various fields, making automation smarter and more reliable.
最新情報をメールで取得
登録
© Couger Inc. All rights reserved.