[{"data":1,"prerenderedAt":10578},["ShallowReactive",2],{"blog-category-recruiting":3},[4,602,719,911,2008,2199,2313,2555,2767,3197,5017,6076,8056,9783,9974,10226,10484],{"_path":5,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":9,"description":10,"draft":7,"publicationDate":11,"updatedAt":11,"image":12,"author":13,"ogTitle":18,"ogDescription":19,"twitterTitle":20,"twitterDescription":21,"keywords":22,"tags":23,"head":30,"category":42,"body":43,"_type":596,"_id":597,"_source":598,"_file":599,"_stem":600,"_extension":601},"/blog/ai-analyst-should-show-its-work","blog",false,"","Your AI Analyst Should Show Its Work: Why AI Analytics Needs Citations","AI analytics tools that don't show their data sources are a governance risk. Learn why AI observability, data citations, and audit trails are essential to prevent hallucinated metrics from reaching leadership decisions.","2026-02-16","/images/blog/anamap-citations-reddit-post.jpg",{"id":14,"name":15,"role":16,"twitter":17},"alex-schlee","Alex Schlee","Product Manager","@alexschlee","Your AI Analyst Should Show Its Work: Prevent Hallucinated Analytics","A company made Q4 decisions on AI-hallucinated data. Here's why AI analytics without citations is a governance failure, and how to fix it.","Your AI Analyst Should Show Its Work","A company discovered their AI had been hallucinating analytics numbers for 3 months. Territory decisions, board decks, all fake. Here's what we learned.","AI analytics hallucination, AI observability, AI data citations, AI governance analytics, AI analyst audit trail, prevent AI hallucinated metrics, AI analytics trust, AI traceability, LLM analytics risk, AI generated insights verification",[24,25,26,27,28,29],"AI governance","AI analytics","data citations","observability","hallucination","analytics trust",{"meta":31},[32,34,37,39],{"name":33,"content":22},"keywords",{"name":35,"content":36},"robots","index, follow",{"name":38,"content":15},"author",{"name":40,"content":41},"copyright","© 2026 Anamap","AI & Analytics",{"type":44,"children":45,"toc":586},"root",[46,54,61,77,82,95,100,105,112,117,125,131,136,166,171,199,204,209,217,223,228,246,251,279,287,292,297,303,308,313,318,390,397,402,408,416,421,426,438,444,449,457,470,478,501,506,511,516,521,555,560,566,571,576,581],{"type":47,"tag":48,"props":49,"children":50},"element","p",{},[51],{"type":52,"value":53},"text","AI analytics tools that generate insights without showing the underlying data source are a governance risk. If your AI analyst cannot show the exact query, raw data, and reasoning behind every number it presents, it is not production-ready. This post explains why AI observability and data citations are essential, and what happens when they're missing.",{"type":47,"tag":55,"props":56,"children":58},"h2",{"id":57},"a-reddit-post-that-should-scare-every-analytics-leader",[59],{"type":52,"value":60},"A Reddit Post That Should Scare Every Analytics Leader",{"type":47,"tag":48,"props":62,"children":63},{},[64,66,75],{"type":52,"value":65},"Last week I saw ",{"type":47,"tag":67,"props":68,"children":72},"a",{"href":69,"rel":70},"https://www.reddit.com/r/analytics/s/dw8ZRqPf5H",[71],"nofollow",[73],{"type":52,"value":74},"a post on Reddit",{"type":52,"value":76}," that made my stomach drop.",{"type":47,"tag":48,"props":78,"children":79},{},[80],{"type":52,"value":81},"A company had been using an AI agent to answer leadership questions about metrics. It was fast. Confident. Detailed. Everyone loved it.",{"type":47,"tag":48,"props":83,"children":84},{},[85,87,93],{"type":52,"value":86},"Three months later, they discovered it had been ",{"type":47,"tag":88,"props":89,"children":90},"strong",{},[91],{"type":52,"value":92},"hallucinating numbers",{"type":52,"value":94},".",{"type":47,"tag":48,"props":96,"children":97},{},[98],{"type":52,"value":99},"Territory decisions were made on data that didn't exist. The board saw fake insights. The AI invented plausible-sounding percentages and nobody questioned it because it sounded right.",{"type":47,"tag":48,"props":101,"children":102},{},[103],{"type":52,"value":104},"Legal got involved. Q4 decisions had to be reviewed. People might get fired.",{"type":47,"tag":106,"props":107,"children":111},"blog-image",{"alt":108,"size":109,"src":12,"float":110},"A Reddit post describing how an AI analytics agent hallucinated data for 3 months, leading to wrong business decisions","small","right",[],{"type":47,"tag":48,"props":113,"children":114},{},[115],{"type":52,"value":116},"Here's the uncomfortable truth:",{"type":47,"tag":48,"props":118,"children":119},{},[120],{"type":47,"tag":88,"props":121,"children":122},{},[123],{"type":52,"value":124},"The AI didn't fail because it was malicious. It failed because it wasn't governed.",{"type":47,"tag":55,"props":126,"children":128},{"id":127},"we-would-never-accept-this-from-a-human-analyst",[129],{"type":52,"value":130},"We Would Never Accept This From a Human Analyst",{"type":47,"tag":48,"props":132,"children":133},{},[134],{"type":52,"value":135},"If an analyst presents insights in a meeting, we expect:",{"type":47,"tag":137,"props":138,"children":139},"ul",{},[140,146,151,156,161],{"type":47,"tag":141,"props":142,"children":143},"li",{},[144],{"type":52,"value":145},"The exact metric definition",{"type":47,"tag":141,"props":147,"children":148},{},[149],{"type":52,"value":150},"The time range",{"type":47,"tag":141,"props":152,"children":153},{},[154],{"type":52,"value":155},"The data source",{"type":47,"tag":141,"props":157,"children":158},{},[159],{"type":52,"value":160},"The breakdown logic",{"type":47,"tag":141,"props":162,"children":163},{},[164],{"type":52,"value":165},"The raw numbers behind the summary",{"type":47,"tag":48,"props":167,"children":168},{},[169],{"type":52,"value":170},"If they say \"conversion increased 12%,\" someone can ask:",{"type":47,"tag":137,"props":172,"children":173},{},[174,183,191],{"type":47,"tag":141,"props":175,"children":176},{},[177],{"type":47,"tag":178,"props":179,"children":180},"em",{},[181],{"type":52,"value":182},"Compared to what?",{"type":47,"tag":141,"props":184,"children":185},{},[186],{"type":47,"tag":178,"props":187,"children":188},{},[189],{"type":52,"value":190},"Over what time period?",{"type":47,"tag":141,"props":192,"children":193},{},[194],{"type":47,"tag":178,"props":195,"children":196},{},[197],{"type":52,"value":198},"Based on which events?",{"type":47,"tag":48,"props":200,"children":201},{},[202],{"type":52,"value":203},"And they need to answer.",{"type":47,"tag":48,"props":205,"children":206},{},[207],{"type":52,"value":208},"We don't allow humans to say, \"Trust me.\"",{"type":47,"tag":48,"props":210,"children":211},{},[212],{"type":47,"tag":88,"props":213,"children":214},{},[215],{"type":52,"value":216},"Why are we letting AI do that?",{"type":47,"tag":55,"props":218,"children":220},{"id":219},"the-real-risk-no-evidence-behind-the-answer",[221],{"type":52,"value":222},"The Real Risk: No Evidence Behind the Answer",{"type":47,"tag":48,"props":224,"children":225},{},[226],{"type":52,"value":227},"Most AI analytics tools generate:",{"type":47,"tag":137,"props":229,"children":230},{},[231,236,241],{"type":47,"tag":141,"props":232,"children":233},{},[234],{"type":52,"value":235},"Executive summaries",{"type":47,"tag":141,"props":237,"children":238},{},[239],{"type":52,"value":240},"Polished explanations",{"type":47,"tag":141,"props":242,"children":243},{},[244],{"type":52,"value":245},"Confident recommendations",{"type":47,"tag":48,"props":247,"children":248},{},[249],{"type":52,"value":250},"But they don't expose:",{"type":47,"tag":137,"props":252,"children":253},{},[254,259,264,269,274],{"type":47,"tag":141,"props":255,"children":256},{},[257],{"type":52,"value":258},"The exact API query that was run",{"type":47,"tag":141,"props":260,"children":261},{},[262],{"type":52,"value":263},"The dimensions and metrics used",{"type":47,"tag":141,"props":265,"children":266},{},[267],{"type":52,"value":268},"The raw rows returned from the data source",{"type":47,"tag":141,"props":270,"children":271},{},[272],{"type":52,"value":273},"The aggregation logic",{"type":47,"tag":141,"props":275,"children":276},{},[277],{"type":52,"value":278},"The reasoning chain behind conclusions",{"type":47,"tag":48,"props":280,"children":281},{},[282],{"type":47,"tag":88,"props":283,"children":284},{},[285],{"type":52,"value":286},"That's where things break.",{"type":47,"tag":48,"props":288,"children":289},{},[290],{"type":52,"value":291},"If the AI fabricates numbers internally, you won't see it. If it mixes time ranges, you won't know. If it combines mismatched dimensions, it looks right but isn't.",{"type":47,"tag":48,"props":293,"children":294},{},[295],{"type":52,"value":296},"Without evidence, you're just taking the AI's word for it.",{"type":47,"tag":55,"props":298,"children":300},{"id":299},"what-we-changed-in-anamap",[301],{"type":52,"value":302},"What We Changed in Anamap",{"type":47,"tag":48,"props":304,"children":305},{},[306],{"type":52,"value":307},"After seeing that Reddit post, I doubled down on something I already believed: AI analytics must show its work.",{"type":47,"tag":48,"props":309,"children":310},{},[311],{"type":52,"value":312},"So we built explicit data citations into every insight and recommendation Anamap produces.",{"type":47,"tag":48,"props":314,"children":315},{},[316],{"type":52,"value":317},"Now when Cartos AI produces an insight, you can:",{"type":47,"tag":137,"props":319,"children":320},{},[321,340,350,360,370,380],{"type":47,"tag":141,"props":322,"children":323},{},[324,329,331,338],{"type":47,"tag":88,"props":325,"children":326},{},[327],{"type":52,"value":328},"See the exact data request",{"type":52,"value":330}," sent to the source (e.g., GA4 ",{"type":47,"tag":332,"props":333,"children":335},"code",{"className":334},[],[336],{"type":52,"value":337},"runReport",{"type":52,"value":339},")",{"type":47,"tag":141,"props":341,"children":342},{},[343,348],{"type":47,"tag":88,"props":344,"children":345},{},[346],{"type":52,"value":347},"Inspect metrics and dimensions",{"type":52,"value":349}," used in the query",{"type":47,"tag":141,"props":351,"children":352},{},[353,358],{"type":47,"tag":88,"props":354,"children":355},{},[356],{"type":52,"value":357},"Verify the date range",{"type":52,"value":359}," the data covers",{"type":47,"tag":141,"props":361,"children":362},{},[363,368],{"type":47,"tag":88,"props":364,"children":365},{},[366],{"type":52,"value":367},"See rows returned",{"type":52,"value":369}," from the API response",{"type":47,"tag":141,"props":371,"children":372},{},[373,378],{"type":47,"tag":88,"props":374,"children":375},{},[376],{"type":52,"value":377},"Review the AI rationale",{"type":52,"value":379}," for why it queried that specific data",{"type":47,"tag":141,"props":381,"children":382},{},[383,388],{"type":47,"tag":88,"props":384,"children":385},{},[386],{"type":52,"value":387},"Validate the raw numbers",{"type":52,"value":389}," behind every claim",{"type":47,"tag":106,"props":391,"children":396},{"alt":392,"size":393,"src":394,"caption":395},"Cartos AI showing insights with numbered citation markers that link to the underlying data source","medium","/images/blog/anamap-citations-1.png","Each numbered citation links directly to the data source that produced the numbers referenced in that insight.",[],{"type":47,"tag":48,"props":398,"children":399},{},[400],{"type":52,"value":401},"Click any citation to see the full query details.",{"type":47,"tag":106,"props":403,"children":407},{"alt":404,"size":393,"src":405,"caption":406},"The Data Sources panel showing the exact Google Analytics query, including query type, date range, metrics, dimensions, rows returned, AI rationale, and the raw data preview","/images/blog/anamap-citations-2.webp","The query type, date range, metrics, dimensions, rows returned, and AI rationale are all visible.",[],{"type":47,"tag":48,"props":409,"children":410},{},[411],{"type":47,"tag":88,"props":412,"children":413},{},[414],{"type":52,"value":415},"The AI does not generate the numbers.",{"type":47,"tag":48,"props":417,"children":418},{},[419],{"type":52,"value":420},"The data comes directly from API calls to your data source (e.g., Google Analytics, Amplitude). The AI analyzes those results, but it cannot fabricate numbers inside citations because they are tied to actual query responses.",{"type":47,"tag":48,"props":422,"children":423},{},[424],{"type":52,"value":425},"Every insight is anchored to retrievable source data.",{"type":47,"tag":48,"props":427,"children":428},{},[429,431,436],{"type":52,"value":430},"If an executive asks, ",{"type":47,"tag":178,"props":432,"children":433},{},[434],{"type":52,"value":435},"\"Where did that number come from?\"",{"type":52,"value":437}," you can click and show them.",{"type":47,"tag":55,"props":439,"children":441},{"id":440},"citations-are-a-requirement-now",[442],{"type":52,"value":443},"Citations Are a Requirement Now",{"type":47,"tag":48,"props":445,"children":446},{},[447],{"type":52,"value":448},"AI in analytics changes the surface area of risk.",{"type":47,"tag":48,"props":450,"children":451},{},[452],{"type":47,"tag":88,"props":453,"children":454},{},[455],{"type":52,"value":456},"Before:",{"type":47,"tag":137,"props":458,"children":459},{},[460,465],{"type":47,"tag":141,"props":461,"children":462},{},[463],{"type":52,"value":464},"Humans could miscalculate",{"type":47,"tag":141,"props":466,"children":467},{},[468],{"type":52,"value":469},"Humans could misunderstand definitions",{"type":47,"tag":48,"props":471,"children":472},{},[473],{"type":47,"tag":88,"props":474,"children":475},{},[476],{"type":52,"value":477},"Now:",{"type":47,"tag":137,"props":479,"children":480},{},[481,486,491,496],{"type":47,"tag":141,"props":482,"children":483},{},[484],{"type":52,"value":485},"AI can misinterpret schema",{"type":47,"tag":141,"props":487,"children":488},{},[489],{"type":52,"value":490},"AI can select the wrong breakdown",{"type":47,"tag":141,"props":492,"children":493},{},[494],{"type":52,"value":495},"AI can summarize incorrectly",{"type":47,"tag":141,"props":497,"children":498},{},[499],{"type":52,"value":500},"AI can hallucinate if not constrained",{"type":47,"tag":48,"props":502,"children":503},{},[504],{"type":52,"value":505},"Governance has to keep up.",{"type":47,"tag":48,"props":507,"children":508},{},[509],{"type":52,"value":510},"Citations are not a design detail. They are a control mechanism.",{"type":47,"tag":48,"props":512,"children":513},{},[514],{"type":52,"value":515},"In finance, we require audit trails. In engineering, we require logs. In analytics, we now need the same thing.",{"type":47,"tag":48,"props":517,"children":518},{},[519],{"type":52,"value":520},"If your AI analyst cannot show:",{"type":47,"tag":522,"props":523,"children":524},"ol",{},[525,535,545],{"type":47,"tag":141,"props":526,"children":527},{},[528,533],{"type":47,"tag":88,"props":529,"children":530},{},[531],{"type":52,"value":532},"The query",{"type":52,"value":534}," it sent to the data source",{"type":47,"tag":141,"props":536,"children":537},{},[538,543],{"type":47,"tag":88,"props":539,"children":540},{},[541],{"type":52,"value":542},"The data",{"type":52,"value":544}," that came back",{"type":47,"tag":141,"props":546,"children":547},{},[548,553],{"type":47,"tag":88,"props":549,"children":550},{},[551],{"type":52,"value":552},"The logic",{"type":52,"value":554}," behind how it interpreted the results",{"type":47,"tag":48,"props":556,"children":557},{},[558],{"type":52,"value":559},"It is not production-ready.",{"type":47,"tag":55,"props":561,"children":563},{"id":562},"the-standard-has-changed",[564],{"type":52,"value":565},"The Standard Has Changed",{"type":47,"tag":48,"props":567,"children":568},{},[569],{"type":52,"value":570},"We've always expected analysts to support claims with data.",{"type":47,"tag":48,"props":572,"children":573},{},[574],{"type":52,"value":575},"We should demand the same from AI.",{"type":47,"tag":48,"props":577,"children":578},{},[579],{"type":52,"value":580},"Anamap is built around observability and governance, not just insight generation. Fast answers are useless if they're wrong.",{"type":47,"tag":48,"props":582,"children":583},{},[584],{"type":52,"value":585},"If AI is going to sit in leadership meetings, it needs to be auditable.",{"title":8,"searchDepth":587,"depth":588,"links":589},2,3,[590,591,592,593,594,595],{"id":57,"depth":587,"text":60},{"id":127,"depth":587,"text":130},{"id":219,"depth":587,"text":222},{"id":299,"depth":587,"text":302},{"id":440,"depth":587,"text":443},{"id":562,"depth":587,"text":565},"markdown","content:blog:ai-analyst-should-show-its-work.md","content","blog/ai-analyst-should-show-its-work.md","blog/ai-analyst-should-show-its-work","md",{"_path":603,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":604,"description":605,"draft":7,"publicationDate":606,"image":607,"author":608,"head":609,"category":617,"body":618,"_type":596,"_id":716,"_source":598,"_file":717,"_stem":718,"_extension":601},"/blog/analytics-maps-turbocharge-insights","Analytics Maps Will Turbocharge Your Insights","Analytics is the way that we measure our businesses but if you don't know which measurements you have available it's hard to maximize your company's ability to drive positive change.","2024-05-27","/images/blog/d27405ee-e839-464b-adb5-48abb6e01055.webp",{"id":14,"name":15,"role":16},{"meta":610},[611,613,614,615],{"name":33,"content":612},"analytics, business strategy, productivity",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},"© 2024 Anamap","Mapping",{"type":44,"children":619,"toc":710},[620,626,631,637,642,665,671,676,699,705],{"type":47,"tag":55,"props":621,"children":623},{"id":622},"data-analytics-knowledge-a-better-way",[624],{"type":52,"value":625},"Data Analytics Knowledge A Better Way",{"type":47,"tag":48,"props":627,"children":628},{},[629],{"type":52,"value":630},"Analytics is the way that we measure the performance of our organizations but if you don't know which measurements you have available it's hard to maximize your organization's ability to drive positive change. Using an analytics map makes it easier for the analysts and stakeholders in your business to understand the story of your data and use those insights to propel your organization to greater heights.",{"type":47,"tag":55,"props":632,"children":634},{"id":633},"analytics-map-benefits",[635],{"type":52,"value":636},"Analytics Map Benefits",{"type":47,"tag":48,"props":638,"children":639},{},[640],{"type":52,"value":641},"Whether you choose to go a more traditional route for your analytics mapping via spreadsheets or you decide to embrace visual analytics mapping with Anamap there are many benefits for everyone who is involved in the data analytics.",{"type":47,"tag":137,"props":643,"children":644},{},[645,650,655,660],{"type":47,"tag":141,"props":646,"children":647},{},[648],{"type":52,"value":649},"Creating a visual analytics map helps to spread the knowledge about what events and attributes (also called properties) are available to use for reporting or exploration. Many organizations have a singular person or singular team that holds all of this knowledge which makes it harder to spread. These visual analytics maps also make your implementation knowledge less prone to tribal knowledge related to key employees leaving.",{"type":47,"tag":141,"props":651,"children":652},{},[653],{"type":52,"value":654},"Maps can also act as a form of analytics contract with your engineering teams which makes it easier to understand when data collection has changed or become broken.",{"type":47,"tag":141,"props":656,"children":657},{},[658],{"type":52,"value":659},"Because there is documentation about which events and attributes are tracked analysts can more easily find additional avenues of investigation and exploration to find stories in your data that haven't been told. Maybe there is an under utilized attribute in your events that shows how two groups of users behave entirely different and it was only discovered thanks to an analyst noticing it in an analytics map.",{"type":47,"tag":141,"props":661,"children":662},{},[663],{"type":52,"value":664},"The discoverability of the data schema means that anyone from across your organization can become an analyst. The knowledge of which events to use or which attributes are available can be shared with everyone. Any stakeholder can look into the data to find meaningful insights which means they aren't taxing the analytics team, who frequently ends up as a bottleneck. This additional free time for your analytics team means they are able to spend more time on deeper analysis and less time on reporting which means more insights and more productivity for your organization or business.",{"type":47,"tag":55,"props":666,"children":668},{"id":667},"visual-analytics-map-anamap-benefits",[669],{"type":52,"value":670},"Visual Analytics Map (Anamap) Benefits",{"type":47,"tag":48,"props":672,"children":673},{},[674],{"type":52,"value":675},"Historically, businesses relied on spreadsheets to keep track of what data was being collected, on which pages of their site, and where that data went. These spreadsheets were cumbersome both for the people maintaining them and also for the end users trying to decipher them. Anamap was created with the core concept that making your organization's analytics implementation easier to understand would mean faster insights with fewer stumbling blocks. Here are some reasons using a visual analytics map is better than a standard spreadsheet-based solution design document (or equivalent).",{"type":47,"tag":137,"props":677,"children":678},{},[679,684,689,694],{"type":47,"tag":141,"props":680,"children":681},{},[682],{"type":52,"value":683},"It's visual. This might seem obvious but providing users something that's more visually organized makes it easier for them to userstand what they are seeing and make use of it. Reading a spreadsheet map of data analytics implementation can make someone's eyes glaze over. Most spreadsheets look the same and it's difficult to find the page or event you're interested in. Making the map visual means it's easy to identify pages and events quickly, see how they are interconnected, and understand what attributes are being collected.",{"type":47,"tag":141,"props":685,"children":686},{},[687],{"type":52,"value":688},"It's relationship driven (in a database). All the attributes, events, and views (pages) are centralized so that changes do not need to be copied multiple times across a document when one thing is changed. This is a common pitfall of traditional methods of keeping track of this information and results in documents becoming more and more inaccurate over time. The relational structure of Anamap means people will need to spend less time on maintennance and creates a cleaner data environment.",{"type":47,"tag":141,"props":690,"children":691},{},[692],{"type":52,"value":693},"It's packed with tools to make building maps easier. Tools like the built-in site crawler and live screenshot feature make it easy to get started creating views for your company and to make sure the image representing your page is always up-to-date.",{"type":47,"tag":141,"props":695,"children":696},{},[697],{"type":52,"value":698},"There's growth pontential. Having all your data mapped in one place (along with the relationships between everything) means the data schema can be used for many different things. In the future, Anamap may be able to connect with your Customer Data Platform (CDP) of your Analytics Platform to help manage your data plan. It could also be connected with an AI agent to help answer any questions users may have where they'd prefer a more conversational way to get the information.",{"type":47,"tag":55,"props":700,"children":702},{"id":701},"should-my-company-create-an-analytics-map",[703],{"type":52,"value":704},"Should My Company Create An Analytics Map?",{"type":47,"tag":48,"props":706,"children":707},{},[708],{"type":52,"value":709},"Yes, 100% without question. No matter which analytics platform you're using like Google Analytics, Adobe Analytics, Amplitude, Mixpanel, or an entirely different data collection system mapping your analytics is a worthy endeavor. The efficiencies it provides in gaining those valuable insights can forever change the trajectory of your organization.",{"title":8,"searchDepth":587,"depth":588,"links":711},[712,713,714,715],{"id":622,"depth":587,"text":625},{"id":633,"depth":587,"text":636},{"id":667,"depth":587,"text":670},{"id":701,"depth":587,"text":704},"content:blog:analytics-maps-turbocharge-insights.md","blog/analytics-maps-turbocharge-insights.md","blog/analytics-maps-turbocharge-insights",{"_path":720,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":721,"description":722,"draft":7,"publicationDate":723,"image":724,"author":725,"head":726,"category":733,"body":734,"_type":596,"_id":908,"_source":598,"_file":909,"_stem":910,"_extension":601},"/blog/being-a-data-driven-organization","What Does it Look Like to Be a Data-Driven Organization?","Is your business hitting all the key elements of becoming a data-driven organization? Focusing on strategic data use, self-service reporting, methodical testing, and bringing data to the table.","2024-09-19","/images/blog/7af015cb-547c-4b94-83da-e706bca4cafa.webp",{"id":14,"name":15,"role":16},{"meta":727},[728,730,731,732],{"name":33,"content":729},"data, data driven, strategic data collection, self-service reporting, data democratization, data discoverability",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},"Business Insight",{"type":44,"children":735,"toc":899},[736,742,754,763,768,773,779,784,789,794,800,812,817,822,827,833,838,843,849,854,862,867,873,878,883,888,894],{"type":47,"tag":55,"props":737,"children":739},{"id":738},"intro",[740],{"type":52,"value":741},"Intro",{"type":47,"tag":48,"props":743,"children":744},{},[745,747,752],{"type":52,"value":746},"Beyond being a buzzword in business, the concept of being data-driven might be one of the single most important aspects of running a successful business. If you have read many business strategy books you'll have undoubtedly have seen an adage like, \"If you can't measure it, you can't make a decision based on it\". Eric Ries, in ",{"type":47,"tag":178,"props":748,"children":749},{},[750],{"type":52,"value":751},"The Lean Startup",{"type":52,"value":753},", emphasizes that data-driven decision making is essential to success:",{"type":47,"tag":755,"props":756,"children":757},"blockquote",{},[758],{"type":47,"tag":48,"props":759,"children":760},{},[761],{"type":52,"value":762},"\"A startup should measure everything that matters, and nothing that doesn’t. This means focusing on metrics that will help you understand whether your product is growing, whether customers are happy, and whether you have a sustainable business model.\"",{"type":47,"tag":48,"props":764,"children":765},{},[766],{"type":52,"value":767},"Ries dedicates multiple chapters of the book to the discussion of data collection and testing. Ultimately, he points out that without data to measure the business and guide experimentation, businesses are essentially operating in the dark. Despite the focus on startups the lessons about the use of data and testing can be applied to any mature business as well.",{"type":47,"tag":48,"props":769,"children":770},{},[771],{"type":52,"value":772},"Making decisions based on gut instincts is a necessary part guiding a company but using your gut exclusively will hamper the success of your business. The decisions your company makes should be based on data and testing as often as is possible.",{"type":47,"tag":55,"props":774,"children":776},{"id":775},"strategic-data-collection",[777],{"type":52,"value":778},"Strategic Data Collection",{"type":47,"tag":48,"props":780,"children":781},{},[782],{"type":52,"value":783},"Some analysts might tell you to collect any events and attributes you possibly can. On the surface, this seems like sound advice. More data is better right? The reality is that this approach is rarely feasible due to resource constraints. Developer time is at a premium in any organization and analytics vendor costs can quickly get out of control if your company tries to track every user interaction. The key is to be strategic about which events you collect and ensure that you’re gathering data that will drive meaningful insights.",{"type":47,"tag":48,"props":785,"children":786},{},[787],{"type":52,"value":788},"One of the first questions I like to ask stakeholders when they say they need some specific piece of data is: “what decisions would you be able to make if you had this data that you can’t make right now?”. Whether they have an answer or not it at least challenges them to think about the application of the data and not just collecting it “just in case someone asks for it”. What do your users really care about? What behaviors predict churn, conversions, or high engagement? Once you have laid out these key behaviors you can begin designing your data collection strategy.",{"type":47,"tag":48,"props":790,"children":791},{},[792],{"type":52,"value":793},"In a separate post I’ll lay out some data collection strategies that will help level up your organization’s data.",{"type":47,"tag":55,"props":795,"children":797},{"id":796},"reporting-tools-for-everyone",[798],{"type":52,"value":799},"Reporting Tools For Everyone",{"type":47,"tag":48,"props":801,"children":802},{},[803,805,810],{"type":52,"value":804},"Data democratization is the concept that everyone within an organization can get access to the data they need to do their jobs and create their own insight. This widespread availability of the data is critical for making the transition to a data-driven company. Don’t be fooled into thinking that giving everyone access to a dashboard or a few reports is going to be enough. Access ",{"type":47,"tag":88,"props":806,"children":807},{},[808],{"type":52,"value":809},"and",{"type":52,"value":811}," training are the two pillars of a successful democratization strategy. To build a truly data-driven organization, your team must have the correct tools and the knowledge to use those tools effectively.",{"type":47,"tag":48,"props":813,"children":814},{},[815],{"type":52,"value":816},"Everyone within your company, not just the analysts, should be able to leverage your main reporting tools to help them make informed, data-backed decisions. All departments from marketing and product teams to sales and customer support should have access to data that is relevant to their day to day operations. In light of that, choosing a reporting tool that is intuitive for non-technical users but has the flexibility that technical users need for deep querying is essential.",{"type":47,"tag":48,"props":818,"children":819},{},[820],{"type":52,"value":821},"Access alone isn’t enough. Without proper training, employees will inevitably experience friction using the tools or draw invalid conclusions which will discourage further use and completely kill any momentum your company had towards your transition to being data-driven. In addition to the obvious topic of tool use the training should also cover how to interpret that data in a meaningful way. Employees should feel empowered to analyze the data, spot trends, and make informed decisions that align with the company’s strategic goals.",{"type":47,"tag":48,"props":823,"children":824},{},[825],{"type":52,"value":826},"Democratization of your company’s data will lead to better faster decision-making for the entire organization and fosters a culture of transparency and accountability.",{"type":47,"tag":55,"props":828,"children":830},{"id":829},"data-discoverability",[831],{"type":52,"value":832},"Data Discoverability",{"type":47,"tag":48,"props":834,"children":835},{},[836],{"type":52,"value":837},"Self-service reporting enabled by the excellent reporting tool choice and companion training is only possible when the data is easy to understand. Even with the best tools, employees can find it challenging to successfully analyze the data if they don’t understand how the data is structured. Your organization needs to ensure that the data itself is accessible but also that the definitions for which events and attributes are tracked is accessible and easy to interpret. Most organizations utilize a wiki page or a spreadsheet to keep track of their data schema but both of these can be cumbersome.",{"type":47,"tag":48,"props":839,"children":840},{},[841],{"type":52,"value":842},"Anamp is a great option for making your data more discoverable and improving your documentation. By offering visual data maps, clear schema representations, and easy search functionality Anamap makes it easier for anyone in your organization to effectively utilize your data.",{"type":47,"tag":55,"props":844,"children":846},{"id":845},"the-importance-of-testing",[847],{"type":52,"value":848},"The Importance of Testing",{"type":47,"tag":48,"props":850,"children":851},{},[852],{"type":52,"value":853},"A solid testing program in your organization can not only optimize various parts of your product or service but it can also lead to some unexpected insights about your customers. In an ideal world, testing is so interconnected with your company culture that it’s easy (and expected) to test nearly every new product feature, process, marketing campaign, etc. Even complicated elements of your business such as pricing can be successfully tested to find what drives the most customer growth.",{"type":47,"tag":48,"props":855,"children":856},{},[857],{"type":47,"tag":178,"props":858,"children":859},{},[860],{"type":52,"value":861},"Random interjection here: I hear the marketers reading this screaming at me about brand building and slow burn effects like awareness. I concede there are some things that are very challenging to measure, especially in the short timescales typically allowed for testing so those things get a pass. However, just because they are difficult to A/B test or measure doesn’t mean it’s not worth trying.",{"type":47,"tag":48,"props":863,"children":864},{},[865],{"type":52,"value":866},"Testing fosters a mindset where assumptions and opinions are challenged, hypotheses about your customers are formed, and data provides the final verdict. Without methodical testing businesses are essentially throwing pasta at a wall in the dark trying to figure out what stuck. Each experiment (AKA test) provides valuable insights that either confirm your hypotheses or reject them which can ultimately lead your organization down a new path of discovery and understanding; you won’t just know WHAT works better you will begin to understand WHY it works better too.",{"type":47,"tag":55,"props":868,"children":870},{"id":869},"bringing-data-to-the-table",[871],{"type":52,"value":872},"Bringing Data to the Table",{"type":47,"tag":48,"props":874,"children":875},{},[876],{"type":52,"value":877},"In meetings, decisions are often made based on opinions, past experiences, or intuition. These perspectives aren’t without value in a data-driven organization, but decisions should be backed by evidence whenever possible. Bringing data to the table doesn’t mean eliminating opinions, it just means those opinions should be supported with hard facts and insights. When stakeholders come to the table with statements supported by data the discussions are more productive and less frequently devolve into arguments based on pure speculation. Instead of debating subjective viewpoints, teams can focus on interpreting the data and understanding what it conveys about the business problem.",{"type":47,"tag":48,"props":879,"children":880},{},[881],{"type":52,"value":882},"For example, let’s say you’re in a product meeting, and someone suggests a feature change to improve user engagement. Rather than simply moving forward based on intuition, a data-driven approach would involve gathering data from past user behavior, analyzing patterns, and even running small-scale tests to validate the assumption. This process not only reduces risk but also builds confidence that decisions are based on reliable information.",{"type":47,"tag":48,"props":884,"children":885},{},[886],{"type":52,"value":887},"Bringing data to the table can become a powerful tool for alignment and keeps everyone on the same page with a clear understanding of why specific decisions are made.",{"type":47,"tag":55,"props":889,"children":891},{"id":890},"summary",[892],{"type":52,"value":893},"Summary",{"type":47,"tag":48,"props":895,"children":896},{},[897],{"type":52,"value":898},"Being a data-driven organization means making decisions grounded in data rather than gut feelings or assumptions. By strategically collecting the right data, providing everyone with access to reporting tools and training, making data easily discoverable, fostering a testing culture, and ensuring meetings are supported with analysis, your company will be well-equipped to succeed in today's fast-paced business environment. It’s not about turning every decision into an equation—it’s about creating a culture where data enhances every decision your business makes.",{"title":8,"searchDepth":587,"depth":588,"links":900},[901,902,903,904,905,906,907],{"id":738,"depth":587,"text":741},{"id":775,"depth":587,"text":778},{"id":796,"depth":587,"text":799},{"id":829,"depth":587,"text":832},{"id":845,"depth":587,"text":848},{"id":869,"depth":587,"text":872},{"id":890,"depth":587,"text":893},"content:blog:being-a-data-driven-organization.md","blog/being-a-data-driven-organization.md","blog/being-a-data-driven-organization",{"_path":912,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":913,"description":914,"draft":7,"publicationDate":915,"updatedAt":916,"image":917,"author":918,"ogTitle":919,"ogDescription":920,"twitterTitle":921,"twitterDescription":922,"keywords":923,"tags":924,"head":934,"category":42,"body":941,"_type":596,"_id":2005,"_source":598,"_file":2006,"_stem":2007,"_extension":601},"/blog/best-llm-for-analytics","The Best LLM for Analytics in 2026 (Tested on Real Data)","We tested 26 AI models across 58 runs on connected Google Analytics data. MiniMax M2.5 is still the best everyday analytics model, while Gemini 3.5 Flash was best for synthetic-data audits.","2026-02-12","2026-06-18","/images/blog/llm-benchmark-leaderboard-hero.svg",{"id":14,"name":15,"role":16,"twitter":17},"The Best LLM for Analytics in 2026 — 26 Models Tested","We benchmarked 26 AI models across 58 runs on connected GA4 data. See the best LLMs by analytics use case.","The Best LLM for Analytics in 2026 (26 Models Tested)","Claude, GPT-5.5, Gemini, MiniMax, Kimi, Grok, Qwen, GLM and more tested on real analytics workflows.","best LLM for analytics, which AI model for analytics, best AI for Google Analytics, LLM for data analysis, AI analytics comparison 2026, GPT-5.5 analytics, Claude Opus analytics, Gemini analytics, MiniMax analytics",[925,25,926,927,928,929,930,931,932,933],"LLM comparison","best LLM","model recommendation","Google Analytics","data analysis","Claude","GPT-5","Gemini","MiniMax",{"meta":935},[936,938,939,940],{"name":33,"content":937},"best LLM for analytics, which AI model for analytics, best AI for Google Analytics, LLM for data analysis, AI analytics comparison 2026, GPT-5.5 analytics, Claude analytics, Gemini analytics",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},{"type":44,"children":942,"toc":1968},[943,949,959,964,1033,1046,1050,1056,1063,1073,1078,1083,1094,1100,1110,1115,1120,1130,1136,1146,1151,1156,1162,1170,1175,1208,1213,1219,1229,1234,1244,1254,1257,1263,1268,1296,1302,1351,1357,1507,1510,1516,1521,1525,1530,1539,1542,1548,1553,1559,1564,1570,1575,1581,1586,1589,1595,1600,1742,1752,1755,1761,1767,1772,1778,1783,1789,1794,1800,1805,1811,1816,1819,1825,1831,1848,1854,1859,1865,1870,1876,1881,1887,1892,1898,1903],{"type":47,"tag":55,"props":944,"children":946},{"id":945},"the-best-llm-for-analytics-our-recommendation",[947],{"type":52,"value":948},"The Best LLM for Analytics: Our Recommendation",{"type":47,"tag":48,"props":950,"children":951},{},[952,957],{"type":47,"tag":88,"props":953,"children":954},{},[955],{"type":52,"value":956},"The best LLM for everyday analytics is still MiniMax M2.5.",{"type":52,"value":958}," It delivered excellent quality across 3 consecutive runs on connected Google Analytics data, costs about $0.02 per query, and was the fastest model in our Round 2 marketing attribution benchmark at 70 seconds average.",{"type":47,"tag":48,"props":960,"children":961},{},[962],{"type":52,"value":963},"But the answer now depends more on the job:",{"type":47,"tag":965,"props":966,"children":969},"callout-box",{"icon":967,"title":968,"type":890},"mdi-lightbulb","Quick Answer",[970],{"type":47,"tag":137,"props":971,"children":972},{},[973,983,993,1003,1013,1023],{"type":47,"tag":141,"props":974,"children":975},{},[976,981],{"type":47,"tag":88,"props":977,"children":978},{},[979],{"type":52,"value":980},"Best everyday analytics model:",{"type":52,"value":982}," MiniMax M2.5 -- cheapest, fast, excellent quality",{"type":47,"tag":141,"props":984,"children":985},{},[986,991],{"type":47,"tag":88,"props":987,"children":988},{},[989],{"type":52,"value":990},"Best synthetic-data audit model:",{"type":52,"value":992}," Gemini 3.5 Flash -- best evidence-backed Round 3 result",{"type":47,"tag":141,"props":994,"children":995},{},[996,1001],{"type":47,"tag":88,"props":997,"children":998},{},[999],{"type":52,"value":1000},"Best premium deep dive:",{"type":52,"value":1002}," Claude Opus 4.8 or Claude Opus 4.6 -- thorough, expensive, useful for high-stakes analysis",{"type":47,"tag":141,"props":1004,"children":1005},{},[1006,1011],{"type":47,"tag":88,"props":1007,"children":1008},{},[1009],{"type":52,"value":1010},"Best fast classifier:",{"type":52,"value":1012}," Grok 4.3 -- extremely fast, but in Round 3 it did not make data requests",{"type":47,"tag":141,"props":1014,"children":1015},{},[1016,1021],{"type":47,"tag":88,"props":1017,"children":1018},{},[1019],{"type":52,"value":1020},"Best budget consistency:",{"type":52,"value":1022}," Kimi K2.5 or MiniMax M3 -- low cost with strong multi-run performance",{"type":47,"tag":141,"props":1024,"children":1025},{},[1026,1031],{"type":47,"tag":88,"props":1027,"children":1028},{},[1029],{"type":52,"value":1030},"Avoid for production analytics:",{"type":52,"value":1032}," Gemini 2.5 Flash Lite and GPT-5 Mini from Round 1 failure modes",{"type":47,"tag":48,"props":1034,"children":1035},{},[1036,1038,1044],{"type":52,"value":1037},"This recommendation is based on ",{"type":47,"tag":67,"props":1039,"children":1041},{"href":1040},"/blog/llm-analytics-benchmark-leaderboard",[1042],{"type":52,"value":1043},"our benchmark of 26 AI models",{"type":52,"value":1045}," across 58 test runs on connected GA4 data.",{"type":47,"tag":1047,"props":1048,"children":1049},"hr",{},[],{"type":47,"tag":55,"props":1051,"children":1053},{"id":1052},"our-top-picks-by-use-case",[1054],{"type":52,"value":1055},"Our Top Picks by Use Case",{"type":47,"tag":1057,"props":1058,"children":1060},"h3",{"id":1059},"best-for-daily-marketing-analytics",[1061],{"type":52,"value":1062},"Best for Daily Marketing Analytics",{"type":47,"tag":48,"props":1064,"children":1065},{},[1066,1071],{"type":47,"tag":88,"props":1067,"children":1068},{},[1069],{"type":52,"value":1070},"MiniMax M2.5",{"type":52,"value":1072}," -- $0.02/query | 70s avg | 100/100 accuracy",{"type":47,"tag":48,"props":1074,"children":1075},{},[1076],{"type":52,"value":1077},"If you're running analytics queries every day -- checking campaign performance, monitoring conversion rates, investigating traffic patterns -- MiniMax M2.5 is still the clear winner. It delivered excellent results in all 3 Round 2 runs, immediately identified broken attribution tracking, and pivoted to actionable conversion analysis.",{"type":47,"tag":48,"props":1079,"children":1080},{},[1081],{"type":52,"value":1082},"At $0.0003 per 1,000 tokens, it changes the economics of AI-powered analytics. You can run hundreds of routine queries for the cost of a single premium-model deep dive.",{"type":47,"tag":48,"props":1084,"children":1085},{},[1086,1088],{"type":52,"value":1087},"→ ",{"type":47,"tag":67,"props":1089,"children":1091},{"href":1090},"/blog/llm-analytics-benchmark-round-2-consistency#1-minimax-m25",[1092],{"type":52,"value":1093},"See MiniMax M2.5's full benchmark results",{"type":47,"tag":1057,"props":1095,"children":1097},{"id":1096},"best-for-synthetic-data-and-data-quality-audits",[1098],{"type":52,"value":1099},"Best for Synthetic-Data and Data-Quality Audits",{"type":47,"tag":48,"props":1101,"children":1102},{},[1103,1108],{"type":47,"tag":88,"props":1104,"children":1105},{},[1106],{"type":52,"value":1107},"Gemini 3.5 Flash",{"type":52,"value":1109}," -- $0.23/query | 53s avg | 100/100 accuracy",{"type":47,"tag":48,"props":1111,"children":1112},{},[1113],{"type":52,"value":1114},"Round 3 asked a different question: can the model determine whether connected analytics data is real production data or synthetic demo data?",{"type":47,"tag":48,"props":1116,"children":1117},{},[1118],{"type":52,"value":1119},"Gemini 3.5 Flash gave the best evidence-backed answer. It made data requests, kept perfect GA4 field accuracy, and explained synthetic tells like missing traffic attribution, overly tidy event distributions, weak seasonality, and documentation-shaped values.",{"type":47,"tag":48,"props":1121,"children":1122},{},[1123,1124],{"type":52,"value":1087},{"type":47,"tag":67,"props":1125,"children":1127},{"href":1126},"/blog/llm-analytics-benchmark-round-3-synthetic-data#1-gemini-35-flash",[1128],{"type":52,"value":1129},"See the Round 3 synthetic-data benchmark",{"type":47,"tag":1057,"props":1131,"children":1133},{"id":1132},"best-for-strategic-deep-dive-analysis",[1134],{"type":52,"value":1135},"Best for Strategic Deep-Dive Analysis",{"type":47,"tag":48,"props":1137,"children":1138},{},[1139,1144],{"type":47,"tag":88,"props":1140,"children":1141},{},[1142],{"type":52,"value":1143},"Claude Opus 4.8 / Claude Opus 4.6",{"type":52,"value":1145}," -- $0.81-$1.35+ per query | thorough | premium",{"type":47,"tag":48,"props":1147,"children":1148},{},[1149],{"type":52,"value":1150},"For high-stakes strategy work -- quarterly reviews, board narratives, major tracking investigations -- Claude remains one of the strongest choices. Claude Opus 4.6 was the most thorough Round 2 marketing attribution model. Claude Opus 4.8 gave one of the clearest structural synthetic-data explanations in Round 3.",{"type":47,"tag":48,"props":1152,"children":1153},{},[1154],{"type":52,"value":1155},"The tradeoff is cost. Claude Opus 4.8 Fast cost $1.64 in Round 3, and GPT-5.5 cost $1.45. These are not casual daily-query models unless your budget is comfortable.",{"type":47,"tag":1057,"props":1157,"children":1159},{"id":1158},"best-budget-option",[1160],{"type":52,"value":1161},"Best Budget Option",{"type":47,"tag":48,"props":1163,"children":1164},{},[1165],{"type":47,"tag":88,"props":1166,"children":1167},{},[1168],{"type":52,"value":1169},"Grok 4.1 Fast / Gemini 3.1 Flash Lite / MiniMax M3",{"type":47,"tag":48,"props":1171,"children":1172},{},[1173],{"type":52,"value":1174},"Budget recommendations depend on the task:",{"type":47,"tag":137,"props":1176,"children":1177},{},[1178,1188,1198],{"type":47,"tag":141,"props":1179,"children":1180},{},[1181,1186],{"type":47,"tag":88,"props":1182,"children":1183},{},[1184],{"type":52,"value":1185},"Grok 4.1 Fast",{"type":52,"value":1187}," was the best low-cost Round 1 broken-data model.",{"type":47,"tag":141,"props":1189,"children":1190},{},[1191,1196],{"type":47,"tag":88,"props":1192,"children":1193},{},[1194],{"type":52,"value":1195},"Gemini 3.1 Flash Lite",{"type":52,"value":1197}," was the cheapest successful Round 3 model at $0.027, but its schema score was lower.",{"type":47,"tag":141,"props":1199,"children":1200},{},[1201,1206],{"type":47,"tag":88,"props":1202,"children":1203},{},[1204],{"type":52,"value":1205},"MiniMax M3",{"type":52,"value":1207}," gave a nuanced Round 3 synthetic-data answer for $0.055, but it was slow.",{"type":47,"tag":48,"props":1209,"children":1210},{},[1211],{"type":52,"value":1212},"Cheap can be excellent. Cheap can also be dangerous. That is why we benchmark.",{"type":47,"tag":1057,"props":1214,"children":1216},{"id":1215},"best-for-consistency-critical-workflows",[1217],{"type":52,"value":1218},"Best for Consistency-Critical Workflows",{"type":47,"tag":48,"props":1220,"children":1221},{},[1222,1227],{"type":47,"tag":88,"props":1223,"children":1224},{},[1225],{"type":52,"value":1226},"Kimi K2.5",{"type":52,"value":1228}," -- $0.02/query | 125s avg | 100/100 accuracy",{"type":47,"tag":48,"props":1230,"children":1231},{},[1232],{"type":52,"value":1233},"If you're building automated analytics workflows where consistency matters, Kimi K2.5 still stands out. In Round 2, it had the lowest time variance of any model we tested and delivered excellent quality in all 3 runs.",{"type":47,"tag":48,"props":1235,"children":1236},{},[1237,1238],{"type":52,"value":1087},{"type":47,"tag":67,"props":1239,"children":1241},{"href":1240},"/blog/llm-analytics-benchmark-round-2-consistency#2-kimi-k25",[1242],{"type":52,"value":1243},"See Kimi K2.5's full benchmark results",{"type":47,"tag":1245,"props":1246,"children":1253},"inline-cta",{"heading":1247,"icon":1248,"primary-link":1249,"primary-text":1250,"secondary-link":1040,"secondary-text":1251,"subheading":1252},"Skip the benchmarking -- use Anamap","mdi-robot-outline","/register","Try Anamap Free","View Full Leaderboard","Anamap uses rigorously tested AI models to deliver reliable analytics insights out of the box.",[],{"type":47,"tag":1047,"props":1255,"children":1256},{},[],{"type":47,"tag":55,"props":1258,"children":1260},{"id":1259},"how-we-tested-real-data-real-problems",[1261],{"type":52,"value":1262},"How We Tested: Real Data, Real Problems",{"type":47,"tag":48,"props":1264,"children":1265},{},[1266],{"type":52,"value":1267},"Most LLM comparisons test coding puzzles or trivia. We test analytics work:",{"type":47,"tag":137,"props":1269,"children":1270},{},[1271,1276,1281,1286,1291],{"type":47,"tag":141,"props":1272,"children":1273},{},[1274],{"type":52,"value":1275},"Can the model query connected analytics data?",{"type":47,"tag":141,"props":1277,"children":1278},{},[1279],{"type":52,"value":1280},"Can it detect broken tracking?",{"type":47,"tag":141,"props":1282,"children":1283},{},[1284],{"type":52,"value":1285},"Can it avoid fabricating insights?",{"type":47,"tag":141,"props":1287,"children":1288},{},[1289],{"type":52,"value":1290},"Can it explain the evidence behind a conclusion?",{"type":47,"tag":141,"props":1292,"children":1293},{},[1294],{"type":52,"value":1295},"Can it do the same thing reliably across multiple runs?",{"type":47,"tag":1057,"props":1297,"children":1299},{"id":1298},"three-rounds-of-testing",[1300],{"type":52,"value":1301},"Three Rounds of Testing",{"type":47,"tag":137,"props":1303,"children":1304},{},[1305,1321,1336],{"type":47,"tag":141,"props":1306,"children":1307},{},[1308,1319],{"type":47,"tag":88,"props":1309,"children":1310},{},[1311,1317],{"type":47,"tag":67,"props":1312,"children":1314},{"href":1313},"/blog/llm-analytics-benchmark-broken-data",[1315],{"type":52,"value":1316},"Round 1",{"type":52,"value":1318},":",{"type":52,"value":1320}," 10 established models, 1 run each, broken attribution data",{"type":47,"tag":141,"props":1322,"children":1323},{},[1324,1334],{"type":47,"tag":88,"props":1325,"children":1326},{},[1327,1333],{"type":47,"tag":67,"props":1328,"children":1330},{"href":1329},"/blog/llm-analytics-benchmark-round-2-consistency",[1331],{"type":52,"value":1332},"Round 2",{"type":52,"value":1318},{"type":52,"value":1335}," 6 newer models, 3 runs each, marketing attribution consistency",{"type":47,"tag":141,"props":1337,"children":1338},{},[1339,1349],{"type":47,"tag":88,"props":1340,"children":1341},{},[1342,1348],{"type":47,"tag":67,"props":1343,"children":1345},{"href":1344},"/blog/llm-analytics-benchmark-round-3-synthetic-data",[1346],{"type":52,"value":1347},"Round 3",{"type":52,"value":1318},{"type":52,"value":1350}," 10 newer models, 3 runs each, real-vs-synthetic data detection",{"type":47,"tag":1057,"props":1352,"children":1354},{"id":1353},"what-we-measured",[1355],{"type":52,"value":1356},"What We Measured",{"type":47,"tag":1358,"props":1359,"children":1361},"div",{"style":1360},"overflow-x: auto;",[1362],{"type":47,"tag":1363,"props":1364,"children":1365},"table",{},[1366,1390],{"type":47,"tag":1367,"props":1368,"children":1369},"thead",{},[1370],{"type":47,"tag":1371,"props":1372,"children":1373},"tr",{},[1374,1385],{"type":47,"tag":1375,"props":1376,"children":1377},"th",{},[1378],{"type":47,"tag":1379,"props":1380,"children":1382},"span",{"style":1381},"display: inline-block; min-width: 120px;",[1383],{"type":52,"value":1384},"Criteria",{"type":47,"tag":1375,"props":1386,"children":1387},{},[1388],{"type":52,"value":1389},"What It Means",{"type":47,"tag":1391,"props":1392,"children":1393},"tbody",{},[1394,1411,1427,1443,1459,1475,1491],{"type":47,"tag":1371,"props":1395,"children":1396},{},[1397,1406],{"type":47,"tag":1398,"props":1399,"children":1400},"td",{},[1401],{"type":47,"tag":88,"props":1402,"children":1403},{},[1404],{"type":52,"value":1405},"Quality Rating",{"type":47,"tag":1398,"props":1407,"children":1408},{},[1409],{"type":52,"value":1410},"Did the model deliver useful analysis, not just raw data?",{"type":47,"tag":1371,"props":1412,"children":1413},{},[1414,1422],{"type":47,"tag":1398,"props":1415,"children":1416},{},[1417],{"type":47,"tag":88,"props":1418,"children":1419},{},[1420],{"type":52,"value":1421},"Accuracy Score",{"type":47,"tag":1398,"props":1423,"children":1424},{},[1425],{"type":52,"value":1426},"How accurately the model used valid GA4 dimensions and metrics",{"type":47,"tag":1371,"props":1428,"children":1429},{},[1430,1438],{"type":47,"tag":1398,"props":1431,"children":1432},{},[1433],{"type":47,"tag":88,"props":1434,"children":1435},{},[1436],{"type":52,"value":1437},"Data Quality Detection",{"type":47,"tag":1398,"props":1439,"children":1440},{},[1441],{"type":52,"value":1442},"Did it catch broken attribution, synthetic-data tells, or tracking limitations?",{"type":47,"tag":1371,"props":1444,"children":1445},{},[1446,1454],{"type":47,"tag":1398,"props":1447,"children":1448},{},[1449],{"type":47,"tag":88,"props":1450,"children":1451},{},[1452],{"type":52,"value":1453},"Evidence Quality",{"type":47,"tag":1398,"props":1455,"children":1456},{},[1457],{"type":52,"value":1458},"Did it inspect connected data before making claims?",{"type":47,"tag":1371,"props":1460,"children":1461},{},[1462,1470],{"type":47,"tag":1398,"props":1463,"children":1464},{},[1465],{"type":47,"tag":88,"props":1466,"children":1467},{},[1468],{"type":52,"value":1469},"Speed",{"type":47,"tag":1398,"props":1471,"children":1472},{},[1473],{"type":52,"value":1474},"How long did the full model analysis take?",{"type":47,"tag":1371,"props":1476,"children":1477},{},[1478,1486],{"type":47,"tag":1398,"props":1479,"children":1480},{},[1481],{"type":47,"tag":88,"props":1482,"children":1483},{},[1484],{"type":52,"value":1485},"Cost",{"type":47,"tag":1398,"props":1487,"children":1488},{},[1489],{"type":52,"value":1490},"Estimated API cost for the query",{"type":47,"tag":1371,"props":1492,"children":1493},{},[1494,1502],{"type":47,"tag":1398,"props":1495,"children":1496},{},[1497],{"type":47,"tag":88,"props":1498,"children":1499},{},[1500],{"type":52,"value":1501},"Consistency",{"type":47,"tag":1398,"props":1503,"children":1504},{},[1505],{"type":52,"value":1506},"Did repeated runs deliver similar quality?",{"type":47,"tag":1047,"props":1508,"children":1509},{},[],{"type":47,"tag":55,"props":1511,"children":1513},{"id":1512},"the-full-results-26-models",[1514],{"type":52,"value":1515},"The Full Results: 26 Models",{"type":47,"tag":48,"props":1517,"children":1518},{},[1519],{"type":52,"value":1520},"The combined leaderboard now includes 26 models across 3 rounds.",{"type":47,"tag":1522,"props":1523,"children":1524},"llm-benchmark-combined-leaderboard",{},[],{"type":47,"tag":48,"props":1526,"children":1527},{},[1528],{"type":52,"value":1529},"Round 3 changed the interpretation of \"best.\" Grok 4.3 was the fastest successful synthetic-data classifier, but it made zero data requests. Gemini 3.5 Flash was a better evidence-backed audit. GPT-5.5 completed the task but cost $1.45 per run and scored 88/100 on GA4 field accuracy due to metric/dimension compatibility issues.",{"type":47,"tag":48,"props":1531,"children":1532},{},[1533,1534],{"type":52,"value":1087},{"type":47,"tag":67,"props":1535,"children":1536},{"href":1040},[1537],{"type":52,"value":1538},"View the full benchmark leaderboard",{"type":47,"tag":1047,"props":1540,"children":1541},{},[],{"type":47,"tag":55,"props":1543,"children":1545},{"id":1544},"models-to-avoid-for-production-analytics",[1546],{"type":52,"value":1547},"Models to Avoid for Production Analytics",{"type":47,"tag":48,"props":1549,"children":1550},{},[1551],{"type":52,"value":1552},"Not every LLM is safe to use for analytics.",{"type":47,"tag":1057,"props":1554,"children":1556},{"id":1555},"gemini-25-flash-lite-fabricated-traffic-data",[1557],{"type":52,"value":1558},"Gemini 2.5 Flash Lite -- Fabricated Traffic Data",{"type":47,"tag":48,"props":1560,"children":1561},{},[1562],{"type":52,"value":1563},"Despite the data showing 100% \"(not set)\" for all traffic sources in Round 1, Gemini 2.5 Flash Lite invented traffic source data and presented it as real. This is the most dangerous failure mode: a confident wrong answer that could lead to misallocated marketing spend.",{"type":47,"tag":1057,"props":1565,"children":1567},{"id":1566},"gpt-5-mini-misleading-framing",[1568],{"type":52,"value":1569},"GPT-5 Mini -- Misleading Framing",{"type":47,"tag":48,"props":1571,"children":1572},{},[1573],{"type":52,"value":1574},"GPT-5 Mini retrieved the data but framed broken \"(not set)\" values as actionable \"direct traffic\" insights. That is subtler than fabrication, but still dangerous.",{"type":47,"tag":1057,"props":1576,"children":1578},{"id":1577},"claude-fable-5-access-disabled",[1579],{"type":52,"value":1580},"Claude Fable 5 -- Access Disabled",{"type":47,"tag":48,"props":1582,"children":1583},{},[1584],{"type":52,"value":1585},"Claude Fable 5 was selected in Round 3 by the newest-model automation, but failed all attempts through OpenRouter after access to Fable 5 was disabled following a U.S. export-control directive. We replaced it with GPT-5.5 for the final Round 3 analysis.",{"type":47,"tag":1047,"props":1587,"children":1588},{},[],{"type":47,"tag":55,"props":1590,"children":1592},{"id":1591},"cost-comparison-is-the-cheapest-model-good-enough",[1593],{"type":52,"value":1594},"Cost Comparison: Is the Cheapest Model Good Enough?",{"type":47,"tag":48,"props":1596,"children":1597},{},[1598],{"type":52,"value":1599},"Sometimes yes. Sometimes absolutely not.",{"type":47,"tag":1358,"props":1601,"children":1602},{"style":1360},[1603],{"type":47,"tag":1363,"props":1604,"children":1605},{},[1606,1635],{"type":47,"tag":1367,"props":1607,"children":1608},{},[1609],{"type":47,"tag":1371,"props":1610,"children":1611},{},[1612,1620,1625,1630],{"type":47,"tag":1375,"props":1613,"children":1614},{},[1615],{"type":47,"tag":1379,"props":1616,"children":1617},{"style":1381},[1618],{"type":52,"value":1619},"Price Tier",{"type":47,"tag":1375,"props":1621,"children":1622},{},[1623],{"type":52,"value":1624},"Models",{"type":47,"tag":1375,"props":1626,"children":1627},{},[1628],{"type":52,"value":1629},"Quality",{"type":47,"tag":1375,"props":1631,"children":1632},{},[1633],{"type":52,"value":1634},"Risk",{"type":47,"tag":1391,"props":1636,"children":1637},{},[1638,1664,1690,1716],{"type":47,"tag":1371,"props":1639,"children":1640},{},[1641,1649,1654,1659],{"type":47,"tag":1398,"props":1642,"children":1643},{},[1644],{"type":47,"tag":88,"props":1645,"children":1646},{},[1647],{"type":52,"value":1648},"Under $0.05",{"type":47,"tag":1398,"props":1650,"children":1651},{},[1652],{"type":52,"value":1653},"Grok 4.1 Fast, Grok 4.3, Gemini 3.1 Flash Lite, DeepSeek V3.2",{"type":47,"tag":1398,"props":1655,"children":1656},{},[1657],{"type":52,"value":1658},"Often excellent",{"type":47,"tag":1398,"props":1660,"children":1661},{},[1662],{"type":52,"value":1663},"Validate evidence quality",{"type":47,"tag":1371,"props":1665,"children":1666},{},[1667,1675,1680,1685],{"type":47,"tag":1398,"props":1668,"children":1669},{},[1670],{"type":47,"tag":88,"props":1671,"children":1672},{},[1673],{"type":52,"value":1674},"Under $0.10",{"type":47,"tag":1398,"props":1676,"children":1677},{},[1678],{"type":52,"value":1679},"MiniMax M2.5, Kimi K2.5, MiniMax M3, Qwen3.7 Plus",{"type":47,"tag":1398,"props":1681,"children":1682},{},[1683],{"type":52,"value":1684},"Strong value",{"type":47,"tag":1398,"props":1686,"children":1687},{},[1688],{"type":52,"value":1689},"Some models are slow or thin",{"type":47,"tag":1371,"props":1691,"children":1692},{},[1693,1701,1706,1711],{"type":47,"tag":1398,"props":1694,"children":1695},{},[1696],{"type":47,"tag":88,"props":1697,"children":1698},{},[1699],{"type":52,"value":1700},"$0.10-$0.50",{"type":47,"tag":1398,"props":1702,"children":1703},{},[1704],{"type":52,"value":1705},"Qwen3.7 Max, Gemini 3.5 Flash, Gemini 2.5 Flash, Qwen3 Max",{"type":47,"tag":1398,"props":1707,"children":1708},{},[1709],{"type":52,"value":1710},"Strong middle tier",{"type":47,"tag":1398,"props":1712,"children":1713},{},[1714],{"type":52,"value":1715},"Best balance for many teams",{"type":47,"tag":1371,"props":1717,"children":1718},{},[1719,1727,1732,1737],{"type":47,"tag":1398,"props":1720,"children":1721},{},[1722],{"type":47,"tag":88,"props":1723,"children":1724},{},[1725],{"type":52,"value":1726},"Over $0.50",{"type":47,"tag":1398,"props":1728,"children":1729},{},[1730],{"type":52,"value":1731},"Claude Opus models, GPT-5.5, Claude Sonnet 4.5",{"type":47,"tag":1398,"props":1733,"children":1734},{},[1735],{"type":52,"value":1736},"Deep analysis",{"type":47,"tag":1398,"props":1738,"children":1739},{},[1740],{"type":52,"value":1741},"Expensive for routine use",{"type":47,"tag":48,"props":1743,"children":1744},{},[1745,1750],{"type":47,"tag":88,"props":1746,"children":1747},{},[1748],{"type":52,"value":1749},"The takeaway:",{"type":52,"value":1751}," Price alone does not predict quality. MiniMax M2.5 and MiniMax M3 were low-cost winners in different tasks. But Round 1 also showed that cheap models can hallucinate. Always benchmark against the failure modes that matter for your business.",{"type":47,"tag":1047,"props":1753,"children":1754},{},[],{"type":47,"tag":55,"props":1756,"children":1758},{"id":1757},"what-makes-an-llm-good-at-analytics",[1759],{"type":52,"value":1760},"What Makes an LLM Good at Analytics?",{"type":47,"tag":1057,"props":1762,"children":1764},{"id":1763},"_1-data-quality-detection",[1765],{"type":52,"value":1766},"1. Data Quality Detection",{"type":47,"tag":48,"props":1768,"children":1769},{},[1770],{"type":52,"value":1771},"When data is broken, synthetic, or incomplete, the model should say so clearly before recommending action.",{"type":47,"tag":1057,"props":1773,"children":1775},{"id":1774},"_2-evidence-seeking-behavior",[1776],{"type":52,"value":1777},"2. Evidence-Seeking Behavior",{"type":47,"tag":48,"props":1779,"children":1780},{},[1781],{"type":52,"value":1782},"The strongest models inspect the connected data before making claims. Round 3 made this especially visible: a fast answer is less valuable if it does not gather evidence.",{"type":47,"tag":1057,"props":1784,"children":1786},{"id":1785},"_3-analytical-judgment",[1787],{"type":52,"value":1788},"3. Analytical Judgment",{"type":47,"tag":48,"props":1790,"children":1791},{},[1792],{"type":52,"value":1793},"Valid GA4 syntax is table stakes. The real question is what the model does with the results. Good models identify limitations, pivot to better evidence, and explain uncertainty.",{"type":47,"tag":1057,"props":1795,"children":1797},{"id":1796},"_4-actionable-recommendations",[1798],{"type":52,"value":1799},"4. Actionable Recommendations",{"type":47,"tag":48,"props":1801,"children":1802},{},[1803],{"type":52,"value":1804},"The best models tell you what to do next: which tracking to fix, which data to collect, which pages to inspect, and which conclusions are not supported.",{"type":47,"tag":1057,"props":1806,"children":1808},{"id":1807},"_5-consistency-across-runs",[1809],{"type":52,"value":1810},"5. Consistency Across Runs",{"type":47,"tag":48,"props":1812,"children":1813},{},[1814],{"type":52,"value":1815},"LLMs are probabilistic. Round 2 and Round 3 both show that quality can stay stable while timing varies dramatically. Plan for latency variance in production workflows.",{"type":47,"tag":1047,"props":1817,"children":1818},{},[],{"type":47,"tag":55,"props":1820,"children":1822},{"id":1821},"frequently-asked-questions",[1823],{"type":52,"value":1824},"Frequently Asked Questions",{"type":47,"tag":1057,"props":1826,"children":1828},{"id":1827},"what-is-the-best-llm-for-google-analytics",[1829],{"type":52,"value":1830},"What is the best LLM for Google Analytics?",{"type":47,"tag":48,"props":1832,"children":1833},{},[1834,1836,1840,1842,1846],{"type":52,"value":1835},"For everyday marketing analytics, ",{"type":47,"tag":88,"props":1837,"children":1838},{},[1839],{"type":52,"value":1070},{"type":52,"value":1841}," is still the best overall choice from our benchmark. For synthetic-data and data-quality audits, ",{"type":47,"tag":88,"props":1843,"children":1844},{},[1845],{"type":52,"value":1107},{"type":52,"value":1847}," produced the best evidence-backed Round 3 result.",{"type":47,"tag":1057,"props":1849,"children":1851},{"id":1850},"which-is-better-for-analytics-chatgpt-or-claude",[1852],{"type":52,"value":1853},"Which is better for analytics: ChatGPT or Claude?",{"type":47,"tag":48,"props":1855,"children":1856},{},[1857],{"type":52,"value":1858},"It depends on the task. Claude Opus models remain strong for deep analysis. GPT-5.5 correctly identified synthetic demo data in Round 3, but it was expensive and had more GA4 field-compatibility issues than the top Round 3 models.",{"type":47,"tag":1057,"props":1860,"children":1862},{"id":1861},"can-i-use-free-ai-for-analytics",[1863],{"type":52,"value":1864},"Can I use free AI for analytics?",{"type":47,"tag":48,"props":1866,"children":1867},{},[1868],{"type":52,"value":1869},"Free tiers can answer general analytics questions, but our benchmark uses API-connected models with data-source access and multi-turn analysis. That workflow usually requires paid API access.",{"type":47,"tag":1057,"props":1871,"children":1873},{"id":1872},"is-it-safe-to-use-ai-for-analytics-decisions",[1874],{"type":52,"value":1875},"Is it safe to use AI for analytics decisions?",{"type":47,"tag":48,"props":1877,"children":1878},{},[1879],{"type":52,"value":1880},"It depends on the model and the task. Most models in our benchmark delivered useful results, but some fabricated data or framed broken data as insight. Validate models against your own edge cases before production use.",{"type":47,"tag":1057,"props":1882,"children":1884},{"id":1883},"how-much-does-ai-analytics-cost",[1885],{"type":52,"value":1886},"How much does AI analytics cost?",{"type":47,"tag":48,"props":1888,"children":1889},{},[1890],{"type":52,"value":1891},"In our benchmarks, successful runs ranged from a few cents to more than $1.50 depending on model and task. Routine analytics can be very cheap with the right model; premium deep dives are still expensive.",{"type":47,"tag":1057,"props":1893,"children":1895},{"id":1894},"should-i-use-chinese-ai-models-for-analytics",[1896],{"type":52,"value":1897},"Should I use Chinese AI models for analytics?",{"type":47,"tag":48,"props":1899,"children":1900},{},[1901],{"type":52,"value":1902},"MiniMax, Kimi, Qwen, and GLM models performed well in several rounds, often at strong prices. Consider your organization's data residency, privacy, and vendor policy requirements before deploying any third-party model.",{"type":47,"tag":1904,"props":1905,"children":1907},"faq-schema",{":faqs":1906},"[{\"question\":\"What is the best LLM for Google Analytics?\",\"answer\":\"For everyday marketing analytics, MiniMax M2.5 is still the best overall choice from our benchmark. For synthetic-data and data-quality audits, Gemini 3.5 Flash produced the best evidence-backed Round 3 result.\"},{\"question\":\"Which is better for analytics: ChatGPT or Claude?\",\"answer\":\"It depends on the task. Claude Opus models remain strong for deep analysis. GPT-5.5 correctly identified synthetic demo data in Round 3, but it was expensive and had more GA4 field-compatibility issues than the top Round 3 models.\"},{\"question\":\"Can I use free AI for analytics?\",\"answer\":\"Free tiers can answer general analytics questions, but our benchmark uses API-connected models with data-source access and multi-turn analysis. That workflow usually requires paid API access.\"},{\"question\":\"Is it safe to use AI for analytics decisions?\",\"answer\":\"It depends on the model and the task. Most models in our benchmark delivered useful results, but some fabricated data or framed broken data as insight. Validate models against your own edge cases before production use.\"},{\"question\":\"How much does AI analytics cost?\",\"answer\":\"In our benchmarks, successful runs ranged from a few cents to more than $1.50 depending on model and task. Routine analytics can be very cheap with the right model; premium deep dives are still expensive.\"},{\"question\":\"Should I use Chinese AI models for analytics?\",\"answer\":\"MiniMax, Kimi, Qwen, and GLM models performed well in several rounds, often at strong prices. Consider your organizations data residency, privacy, and vendor policy requirements before deploying any third-party model.\"}]",[1908,1914],{"type":47,"tag":55,"props":1909,"children":1911},{"id":1910},"related-reading",[1912],{"type":52,"value":1913},"Related Reading",{"type":47,"tag":137,"props":1915,"children":1916},{},[1917,1927,1937,1947,1957],{"type":47,"tag":141,"props":1918,"children":1919},{},[1920,1925],{"type":47,"tag":67,"props":1921,"children":1922},{"href":1040},[1923],{"type":52,"value":1924},"LLM Analytics Benchmark: The Definitive Leaderboard",{"type":52,"value":1926},": Interactive leaderboard with all 26 models, filters, and sorting",{"type":47,"tag":141,"props":1928,"children":1929},{},[1930,1935],{"type":47,"tag":67,"props":1931,"children":1932},{"href":1344},[1933],{"type":52,"value":1934},"Round 3: Can AI Tell If Analytics Data Is Synthetic?",{"type":52,"value":1936},": Testing GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, and more",{"type":47,"tag":141,"props":1938,"children":1939},{},[1940,1945],{"type":47,"tag":67,"props":1941,"children":1942},{"href":1329},[1943],{"type":52,"value":1944},"Round 2: The Consistency Test",{"type":52,"value":1946},": Testing MiniMax, Kimi, Claude Opus 4.6, GLM 5, Qwen3, and Aurora Alpha",{"type":47,"tag":141,"props":1948,"children":1949},{},[1950,1955],{"type":47,"tag":67,"props":1951,"children":1952},{"href":1313},[1953],{"type":52,"value":1954},"Round 1: Broken Analytics Data",{"type":52,"value":1956},": The original benchmark testing Claude, GPT-5, Gemini, Grok, and DeepSeek",{"type":47,"tag":141,"props":1958,"children":1959},{},[1960,1966],{"type":47,"tag":67,"props":1961,"children":1963},{"href":1962},"/blog/common-analytics-mistakes",[1964],{"type":52,"value":1965},"Common Analytics Mistakes and How to Avoid Them",{"type":52,"value":1967},": The data quality issues that trip up both humans and AI",{"title":8,"searchDepth":587,"depth":588,"links":1969},[1970,1971,1978,1982,1983,1988,1989,1996,2004],{"id":945,"depth":587,"text":948},{"id":1052,"depth":587,"text":1055,"children":1972},[1973,1974,1975,1976,1977],{"id":1059,"depth":588,"text":1062},{"id":1096,"depth":588,"text":1099},{"id":1132,"depth":588,"text":1135},{"id":1158,"depth":588,"text":1161},{"id":1215,"depth":588,"text":1218},{"id":1259,"depth":587,"text":1262,"children":1979},[1980,1981],{"id":1298,"depth":588,"text":1301},{"id":1353,"depth":588,"text":1356},{"id":1512,"depth":587,"text":1515},{"id":1544,"depth":587,"text":1547,"children":1984},[1985,1986,1987],{"id":1555,"depth":588,"text":1558},{"id":1566,"depth":588,"text":1569},{"id":1577,"depth":588,"text":1580},{"id":1591,"depth":587,"text":1594},{"id":1757,"depth":587,"text":1760,"children":1990},[1991,1992,1993,1994,1995],{"id":1763,"depth":588,"text":1766},{"id":1774,"depth":588,"text":1777},{"id":1785,"depth":588,"text":1788},{"id":1796,"depth":588,"text":1799},{"id":1807,"depth":588,"text":1810},{"id":1821,"depth":587,"text":1824,"children":1997},[1998,1999,2000,2001,2002,2003],{"id":1827,"depth":588,"text":1830},{"id":1850,"depth":588,"text":1853},{"id":1861,"depth":588,"text":1864},{"id":1872,"depth":588,"text":1875},{"id":1883,"depth":588,"text":1886},{"id":1894,"depth":588,"text":1897},{"id":1910,"depth":587,"text":1913},"content:blog:best-llm-for-analytics.md","blog/best-llm-for-analytics.md","blog/best-llm-for-analytics",{"_path":1962,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":2009,"description":2010,"draft":7,"publicationDate":2011,"image":2012,"author":2013,"head":2014,"category":733,"body":2021,"_type":596,"_id":2196,"_source":598,"_file":2197,"_stem":2198,"_extension":601},"Don't Make These Common Analytics Mistakes","It’s only natural to fall into the occasional pitfall or two. Hopefully, this list helps your company not fall into some common ones related to analytics.","2024-06-05","/images/blog/c663d297-374d-486f-a9df-fcee3681c490.webp",{"id":14,"name":15,"role":16},{"meta":2015},[2016,2018,2019,2020],{"name":33,"content":2017},"analytics, common mistakes, business insight",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},{"type":44,"children":2022,"toc":2189},[2023,2027,2033,2038,2043,2049,2054,2059,2064,2070,2075,2080,2086,2091,2110,2115,2121,2126,2184],{"type":47,"tag":48,"props":2024,"children":2025},{},[2026],{"type":52,"value":2010},{"type":47,"tag":55,"props":2028,"children":2030},{"id":2029},"_1-thinking-that-switching-to-a-new-analytics-platform-can-solve-all-your-organizations-problems",[2031],{"type":52,"value":2032},"1. Thinking that switching to a new analytics platform can solve all your organization’s problems.",{"type":47,"tag":48,"props":2034,"children":2035},{},[2036],{"type":52,"value":2037},"This seems to be the single most common mistake that I've seen nearly every business make during its existence. I'm sure you've experienced this a time or two or maybe you've even been guilty of perpetuating this mistake. This typically takes the form of a business decision maker saying that they aren't getting enough out of their current data analytics system whether that's Google Analytics, Adobe Analytics, Mixpanel, or others. The line of thinking goes something like this, \"if we only switch to this new vendor, it'll solve all our issues with data\". While I'll admit that there is the odd occasion this is true it isn't true for 95% of companies out there.",{"type":47,"tag":48,"props":2039,"children":2040},{},[2041],{"type":52,"value":2042},"Inevitably companies will switch to a new vendor bearing the burden of doing the migration and all the issues associated with it and will still be unhappy with the outcome. The problem is that people often forget that the tool is just one part of the equation. Having the most powerful available tool on the market means almost nothing if no one at your company knows how to use it. Even with a suboptimal tool, having the right people with the right knowledge can result in great insights. Human capital is by far and away the most important part of your organization getting the most out of data. Invest in your existing people with the right training or find and hire the right outside people, it will make a huge difference.",{"type":47,"tag":55,"props":2044,"children":2046},{"id":2045},"_2-trying-to-track-everything",[2047],{"type":52,"value":2048},"2. Trying to track everything.",{"type":47,"tag":48,"props":2050,"children":2051},{},[2052],{"type":52,"value":2053},"This mistake is most commonly made when the analysts in your organization have control over the analytics implementation or data collection strategy. Having sat in the hotseat myself in front of a table of bristling executives I can tell you that it doesn't feel great when one of them asks a question about some element of the user experience where there is no tracking. Instead of being able to provide a \"yes, let me look that up and get back to you\" you must concede that those interactions aren't tracked which means not only is there no historical data to use for trends but also the turnaround time to get the insights is much longer due to the implementation lag time. Most analysts will try to avoid this situation at all costs by just tracking everything \"just in case\".",{"type":47,"tag":48,"props":2055,"children":2056},{},[2057],{"type":52,"value":2058},"Having too many events with too many different attributes can make it more cumbersome to figure out what event you should really be looking at and it makes the task of maintaining documentation considerably harder. If your company is ever hoping to democratize your data and make business stakeholders data driven decision-makers, you're going to have a hard time if your data structure is too complicated.",{"type":47,"tag":48,"props":2060,"children":2061},{},[2062],{"type":52,"value":2063},"The single most important thing to do when thinking about implementing a new event is just to ask, \"what meaningful business decision could be made if I had this data?\". If you, or the person requesting the data, can't answer this question your company probably doesn't need that data. Don't be afraid to ask executives this question too. I've seen many flights of fancy type requests just because the executive was curious, not because they were expecting to make a decision with the information.",{"type":47,"tag":55,"props":2065,"children":2067},{"id":2066},"_3-using-too-many-events-rather-than-count-attributes",[2068],{"type":52,"value":2069},"3. Using too many events rather than count attributes.",{"type":47,"tag":48,"props":2071,"children":2072},{},[2073],{"type":52,"value":2074},"Most analytics vendors charge by event volume; the more volume you want the more your annual contract is going to cost (assuming you have one). In many cases going over your contracted volume amount and into overage territory will cost your company dearly. The overages can be a very expensive lesson in being conservative with your event collection. In some cases, you may need the full event with all its context and attributes but in many cases you just need a count of how many times an event happened without all the additional attribute information.",{"type":47,"tag":48,"props":2076,"children":2077},{},[2078],{"type":52,"value":2079},"Instead of firing an event for every single time a user interaction is performed you can keep a running count of the number and send it as a user is leaving the page. There are several ways this can be done using sendBeacon and watching for page visibility events as well as the document unload event. Because browsers are all maintained by separate groups of developers and different users have different connection qualities there may be sometimes where this event with counts does not get sent. The counts from this event are going to be directionally accurate, which will allow your business to track it over time and understand whether it is being positively or negatively impacted by changes you’re making.",{"type":47,"tag":55,"props":2081,"children":2083},{"id":2082},"_4-not-using-user-vs-event-level-attributes-properly",[2084],{"type":52,"value":2085},"4. Not using User vs Event level attributes properly.",{"type":47,"tag":48,"props":2087,"children":2088},{},[2089],{"type":52,"value":2090},"A common mistake is overloading event-level attributes with user-level information, and it’s one that can significantly hamper your data management and analysis efforts. Imagine you’re trying to sift through mountains of data, only to find the same user information repeated across every single event. Not only does this lead to redundant data, but it also makes the entire process of managing and analyzing data more cumbersome and less efficient. User-level attributes—like age, location, and subscription type—are meant to provide a snapshot of the user’s profile. These details should describe the user as a whole and not be included repeatedly in every event they trigger. Session level attributes like marketing campaign tracking can also be set at the User-level so all events that happen within that session can be attributed correctly to that campaign.",{"type":47,"tag":48,"props":2092,"children":2093},{},[2094,2096,2101,2103,2108],{"type":52,"value":2095},"Regular updates to user-level attributes and timing those updates correctly are crucial to maintain data accuracy and relevance. In the case of campaigns, you typically want to set them at the very beginning of a new session (or if user clicks through from multiple campaigns in a single session) but you also want to make sure they \"expire\" at the beginning of the new session which means setting them back to an empty state. The timing of user level data updates is also important to consider. Take the case of updating a user's subscription type before versus after the subscription update event. If you update the User-level attribute before the event you may not be able to tell as easily what type of subscription the user was upgrading ",{"type":47,"tag":178,"props":2097,"children":2098},{},[2099],{"type":52,"value":2100},"from",{"type":52,"value":2102},". Conversely, if the User-level attribute is important for figuring out what kind of subscription type they upgrade ",{"type":47,"tag":178,"props":2104,"children":2105},{},[2106],{"type":52,"value":2107},"to",{"type":52,"value":2109}," you it makes sense to update the attribute before the event.",{"type":47,"tag":48,"props":2111,"children":2112},{},[2113],{"type":52,"value":2114},"By avoiding the mistake of overloading event-level attributes with user-level information, and by ensuring regular updates to user-level attributes, you can make your data more manageable and your insights more accurate. This practice not only streamlines your data management but also enhances the precision and relevance of your analytics, ultimately leading to more informed and effective business decisions.",{"type":47,"tag":55,"props":2116,"children":2118},{"id":2117},"_5-assuming-your-data-is-going-to-be-perfectly-clean",[2119],{"type":52,"value":2120},"5. Assuming your data is going to be perfectly clean.",{"type":47,"tag":48,"props":2122,"children":2123},{},[2124],{"type":52,"value":2125},"The Internet is basically the Wild West, and your data is lucky to make it all the way to the server in one piece. There are many different causes of weird data that it's hard to enumerate all of them, but I would coarsely bucket them into these categories:",{"type":47,"tag":137,"props":2127,"children":2128},{},[2129,2140,2151,2162,2173],{"type":47,"tag":141,"props":2130,"children":2131},{},[2132,2138],{"type":47,"tag":1057,"props":2133,"children":2135},{"id":2134},"human-data-entry-errors",[2136],{"type":52,"value":2137},"Human Data Entry Errors",{"type":52,"value":2139},"\nAnywhere a human is involved there is the potential for error. This could be products being categorized incorrectly in your product database or typos in your CMS or incorrect strings in your analytics implementation itself. Benign or not, these issues crop up in every company data set I’ve ever seen. Missing data from human error is easy to spot but miscategorized data can be a real pain to sniff out.",{"type":47,"tag":141,"props":2141,"children":2142},{},[2143,2149],{"type":47,"tag":1057,"props":2144,"children":2146},{"id":2145},"bot-traffic",[2147],{"type":52,"value":2148},"Bot Traffic",{"type":52,"value":2150},"\nEver wonder why some users on your site have 300 page views? Usually, it’s because you have a bot generating way more page views than a real user would. “Don’t analytics platforms block bot activity?” Some commonly known bots like Google Crawler do get blocked by many analytics platforms, however, bots that are created by individuals specifically to scrape data from your site fly under the radar since they only show up on your site. There are some bot detection tools that can help but they are locked into an arms race with the creators of the bots who do not want their bots to be discovered so they aren’t consistently good at detecting the bots.",{"type":47,"tag":141,"props":2152,"children":2153},{},[2154,2160],{"type":47,"tag":1057,"props":2155,"children":2157},{"id":2156},"connectivity-issues",[2158],{"type":52,"value":2159},"Connectivity Issues",{"type":52,"value":2161},"\nPeople tend to take a connection to the internet for granted but at any given time there are many users, your customers, who are experiencing some kind of internet connection issue. Whether that’s because there’s a storm causing their home internet to be intermittent or they’re on a mobile device switching between cell towers or something else entirely, there is a good chance a percentage of your users will experience connection issues while on your site or app. We’ve all been there when we go to click a link (possibly generating a Click event) but your internet cuts out before loading the next page. Many people would look at your event stream and think “how is it possible this user has a click event with no page view after that?”. The next time you see that in your company’s data you’ll know.",{"type":47,"tag":141,"props":2163,"children":2164},{},[2165,2171],{"type":47,"tag":1057,"props":2166,"children":2168},{"id":2167},"browser-differences",[2169],{"type":52,"value":2170},"Browser Differences",{"type":52,"value":2172},"\nThis used to be a much bigger issue in the old days when browsers were more dissimilar, but these days the problems can still be observed in the dark corners of your datasets. Javascript execution and subtle differences in the APIs browsers expose to collect data can be a little different and these differences can result in irregular data collection.",{"type":47,"tag":141,"props":2174,"children":2175},{},[2176,2182],{"type":47,"tag":1057,"props":2177,"children":2179},{"id":2178},"caching-stored-version-of-pages",[2180],{"type":52,"value":2181},"Caching (Stored Version of Pages)",{"type":52,"value":2183},"\nAre you still seeing bad data from something you patched 3 weeks ago? Chances are one of your users is viewing a cached (stored) version of your page from before the fix was released. This can impact both apps and websites since users can take forever to upgrade old apps and there are certain ways for users to effectively save a copy of your website and continue to use it after the live version has changed. These caching issues effectively mean that bad or old data can persist for a surprisingly long time after updates have been made.",{"type":47,"tag":48,"props":2185,"children":2186},{},[2187],{"type":52,"value":2188},"The real takeaway here is to not worry about your data being perfectly clean. As I've outlined it is nearly impossible to eliminate all sources of random data errors. Understand that a (hopefully) small amount of your data is going to be weird and do your best to make directional conclusions based on the majority of the data.",{"title":8,"searchDepth":587,"depth":588,"links":2190},[2191,2192,2193,2194,2195],{"id":2029,"depth":587,"text":2032},{"id":2045,"depth":587,"text":2048},{"id":2066,"depth":587,"text":2069},{"id":2082,"depth":587,"text":2085},{"id":2117,"depth":587,"text":2120},"content:blog:common-analytics-mistakes.md","blog/common-analytics-mistakes.md","blog/common-analytics-mistakes",{"_path":2200,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":2201,"description":2202,"draft":7,"publicationDate":2203,"image":2204,"author":2205,"head":2207,"category":2214,"body":2215,"_type":596,"_id":2310,"_source":598,"_file":2311,"_stem":2312,"_extension":601},"/blog/data-planning-better-data-quality","Don't Waste Time, Implement Your Customer Data Platform Correctly The First Time","Setting up your data plan or tracking plan correctly out-of-the-gate is critical to your company's long term success with a customer data platform (CDP).","2024-06-18","/images/blog/c00785fb-9ebe-4b4a-bb07-e6494d8fa3de.webp",{"id":14,"name":15,"role":2206},"Founder",{"meta":2208},[2209,2211,2212,2213],{"name":33,"content":2210},"data plan, tracking plan, customer data platform, data quality",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},"Data Quality",{"type":44,"children":2216,"toc":2302},[2217,2222,2228,2233,2239,2244,2250,2255,2261,2266,2272,2291,2297],{"type":47,"tag":48,"props":2218,"children":2219},{},[2220],{"type":52,"value":2221},"Setting up your data plan or tracking plan correctly out-of-the-gate is critical to your company's long-term success with a customer data platform (CDP). A poorly executed tracking plan will lock you into a way of organizing your data that will make it more difficult to analyze the data and after you've already collected a bunch of data it can be challenging to convince the organization to start over with a new tracking plan that better serves the business.",{"type":47,"tag":55,"props":2223,"children":2225},{"id":2224},"what-is-a-customer-data-platform",[2226],{"type":52,"value":2227},"What is a Customer Data Platform?",{"type":47,"tag":48,"props":2229,"children":2230},{},[2231],{"type":52,"value":2232},"At their core customer data platforms serve to unify analytics about a customer by tying web analytics, purchase data, email vendor events, or any other off-platform events to a single user in the system. The benefit is that you're able to see a user’s interactions across all your brand touch points and use that data to take key marketing actions. Beyond data unification customer data platforms also provide a means of measuring and enforcing data quality. Segment, ActionIQ, Lytics, and Treasure Data are some of the most popular customer data platforms. The space is rapidly expanding with new companies and new features being released nearly weekly. All of these platforms have one thing in common which is their ability to help enforce data policies in the form of tracking plans that should be leveraged to improve your company's data quality.",{"type":47,"tag":55,"props":2234,"children":2236},{"id":2235},"what-is-data-quality",[2237],{"type":52,"value":2238},"What is Data Quality?",{"type":47,"tag":48,"props":2240,"children":2241},{},[2242],{"type":52,"value":2243},"Data Quality is a measure of how consistent your organization's data is. Higher Data Quality means your data has fewer instances where events are missing properties / attributes (or have extra ones they shouldn't) and those attributes have the correct values attached to them. Naturally, low Data Quality means there are more errors and therefore more noise in your data.",{"type":47,"tag":55,"props":2245,"children":2247},{"id":2246},"why-is-data-quality-important",[2248],{"type":52,"value":2249},"Why is Data Quality Important?",{"type":47,"tag":48,"props":2251,"children":2252},{},[2253],{"type":52,"value":2254},"The more noise is in your data because of lower data quality the harder it is to get clear insights. The patterns tend to be less clear or require more caveats when the data quality is low. If the quality drops low enough people searching the data for answers may come up with different answers which inevitably errors confidence in the data. Once your company has lost confidence in your analytics data it slows down the pace of decision making because leaders thoroughly question insights or may even dismiss them purely due to trust in the data. If you ever find yourself in this position it will be a long uphill battle to regain that lost confidence.",{"type":47,"tag":55,"props":2256,"children":2258},{"id":2257},"how-can-tracking-data-plans-in-cdps-improve-data-quality",[2259],{"type":52,"value":2260},"How can tracking / data plans in CDPs improve data quality?",{"type":47,"tag":48,"props":2262,"children":2263},{},[2264],{"type":52,"value":2265},"The tracking plan in your customer data platform, like Segment, is essentially a contract. The contract is between the platforms that serve as your data sources and your CDP. This contract helps ensure that only valid data is collected and stored in the platform. Depending on which customer data platform your organization works with they have different levels of blocking and different bells and whistles to accompany the various kinds of divergence from plan.",{"type":47,"tag":55,"props":2267,"children":2269},{"id":2268},"what-is-the-right-way-to-setup-my-data-plan-in-segment-amplitude-or-another-cdp",[2270],{"type":52,"value":2271},"What is the right way to setup my data plan in Segment, Amplitude, or another CDP?",{"type":47,"tag":48,"props":2273,"children":2274},{},[2275,2277,2282,2284,2289],{"type":52,"value":2276},"First, each page template should have its own view type event name. Second, each view type event should contain a property or attribute called event_type (or similar) and set to “view” so that your organization is able to easily calculate page views for an entire user experience. There are two major reasons to use this format for your schema. ",{"type":47,"tag":88,"props":2278,"children":2279},{},[2280],{"type":52,"value":2281},"One",{"type":52,"value":2283},": the way that the tracking plans are built can only be as specific as the event name is. In other words, if you have a generic page view event you need to add every conceivable attribute on the event for every page view on your site which means fewer of them can be required. If you use page templates to define the event name, then each page template can have its own unique set of attributes which means your tracking plan can be more specific and therefore more effective. ",{"type":47,"tag":88,"props":2285,"children":2286},{},[2287],{"type":52,"value":2288},"Two",{"type":52,"value":2290},": most analytics platforms for visualizing the data are designed around having unique page event names. When looking at the event stream it's easier to see that a user went from the homepage to the catalog page to the product details page instead of having three non-descript page view events in a row that required expanding to see which specific page they were.",{"type":47,"tag":55,"props":2292,"children":2294},{"id":2293},"is-there-an-easy-way-to-maintain-your-data-plan-tracking-plan",[2295],{"type":52,"value":2296},"Is there an easy way to maintain your data plan / tracking plan?",{"type":47,"tag":48,"props":2298,"children":2299},{},[2300],{"type":52,"value":2301},"Yes, Anamap has you covered. Though most CDPs have a UI for managing your plan they are hard to use and aren't tied to any other system. Anamap allows you to easily manage your attributes, events, and views along with their respective relationships which makes the process of managing the data plan easier which means fewer costly errors. Soon we're releasing a feature that allows you to manage your plans entirely in Anamap and sync the updates to your company's customer data platform. You can manage your tracking plan in one place and have it sync everywhere.",{"title":8,"searchDepth":587,"depth":588,"links":2303},[2304,2305,2306,2307,2308,2309],{"id":2224,"depth":587,"text":2227},{"id":2235,"depth":587,"text":2238},{"id":2246,"depth":587,"text":2249},{"id":2257,"depth":587,"text":2260},{"id":2268,"depth":587,"text":2271},{"id":2293,"depth":587,"text":2296},"content:blog:data-planning-better-data-quality.md","blog/data-planning-better-data-quality.md","blog/data-planning-better-data-quality",{"_path":2314,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":2315,"description":2316,"draft":7,"publicationDate":2317,"image":2318,"author":2319,"head":2320,"category":733,"body":2327,"_type":596,"_id":2552,"_source":598,"_file":2553,"_stem":2554,"_extension":601},"/blog/frontend-vs-backend-data-collection","Frontend Data Collection vs Data Warehouse Ingestion","Should you ingest additional data on the front end or in the database? We'll explain the differences and what the strengths and weaknesses of each are.","2024-08-19","/images/blog/580a2213-1513-4bcd-8f1e-bd442e1a9648.webp",{"id":14,"name":15,"role":16},{"meta":2321},[2322,2324,2325,2326],{"name":33,"content":2323},"data, data collection, web analytics, data engineering, data warehouse",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},{"type":44,"children":2328,"toc":2540},[2329,2335,2340,2345,2350,2355,2361,2367,2410,2416,2449,2455,2460,2483,2488,2531,2535],{"type":47,"tag":55,"props":2330,"children":2332},{"id":2331},"what-is-frontend-data-collection-and-what-is-data-warehouse-ingestion",[2333],{"type":52,"value":2334},"What is Frontend Data Collection and what is Data Warehouse Ingestion?",{"type":47,"tag":48,"props":2336,"children":2337},{},[2338],{"type":52,"value":2339},"Put simply, Frontend Data Collection refers to processing and collecting data on the service or client machine where the analytics events are generated. The most common instance of this is the JavaScript that runs for a site’s web analytics. The actual JavaScript itself runs on the device that is accessing the website.",{"type":47,"tag":48,"props":2341,"children":2342},{},[2343],{"type":52,"value":2344},"Data Warehouse Ingestion refers to the process of pulling data from other server locations (typically) such as an S3 bucket, a production database, CMS, etc. There are many examples of this but a fairly common process is to take a separate table of data about your marketing campaigns (such as from Facebook, Google, etc) and pull it into your data warehouse for the purpose of creating cost per acquisition reporting.",{"type":47,"tag":48,"props":2346,"children":2347},{},[2348],{"type":52,"value":2349},"In many cases, data collected during Frontend Data Collection ultimately ends up in your company Data Warehouse. This can often lead to a discussion about whether certain data should be collected on the frontend or whether it should be ingested directly into the Data Warehouse. To anchor the discussion, one such example where this debate comes up is collecting product information on a product details page view. You can collect the information about the product a user has viewed, like the product name, in the page view event itself or you can store the product ID. When the page view event data lands in the Data Warehouse you can use that product ID to ingest the information about the product from your product database.",{"type":47,"tag":48,"props":2351,"children":2352},{},[2353],{"type":52,"value":2354},"I'm going to cover the pros and cons of each option to help you better decide what the correct choice for your company and user case is.",{"type":47,"tag":55,"props":2356,"children":2358},{"id":2357},"frontend-data-collection",[2359],{"type":52,"value":2360},"Frontend Data Collection",{"type":47,"tag":1057,"props":2362,"children":2364},{"id":2363},"what-are-the-benefits",[2365],{"type":52,"value":2366},"What are the benefits?",{"type":47,"tag":137,"props":2368,"children":2369},{},[2370,2380,2390,2400],{"type":47,"tag":141,"props":2371,"children":2372},{},[2373,2378],{"type":47,"tag":88,"props":2374,"children":2375},{},[2376],{"type":52,"value":2377},"Point-in-time data:",{"type":52,"value":2379}," the data collected and attached to the event truly represents what the user saw. If your store team briefly goofed on the title of a product in the store you'll be able to see that in the data.",{"type":47,"tag":141,"props":2381,"children":2382},{},[2383,2388],{"type":47,"tag":88,"props":2384,"children":2385},{},[2386],{"type":52,"value":2387},"Distributed processing:",{"type":52,"value":2389}," all of the data collection code is running on the client machine which means the processing can be distributed amongst many different computers. Each computer only contributes a few milliseconds of processing time but it saves your cloud resources a lot of compute time in the end.",{"type":47,"tag":141,"props":2391,"children":2392},{},[2393,2398],{"type":47,"tag":88,"props":2394,"children":2395},{},[2396],{"type":52,"value":2397},"Downstream availability:",{"type":52,"value":2399}," because frontend data collection is as far upstream as possible it means whatever data you collect should be available to any additional downstream platforms you choose to send the data to whether that's a CDP, CRM, additional analytics platform, or even a data warehouse.",{"type":47,"tag":141,"props":2401,"children":2402},{},[2403,2408],{"type":47,"tag":88,"props":2404,"children":2405},{},[2406],{"type":52,"value":2407},"Simplicity:",{"type":52,"value":2409}," data that can be collected by the frontend is usually more accessible and requires less work to collect.",{"type":47,"tag":1057,"props":2411,"children":2413},{"id":2412},"what-are-the-drawbacks",[2414],{"type":52,"value":2415},"What are the drawbacks?",{"type":47,"tag":137,"props":2417,"children":2418},{},[2419,2429,2439],{"type":47,"tag":141,"props":2420,"children":2421},{},[2422,2427],{"type":47,"tag":88,"props":2423,"children":2424},{},[2425],{"type":52,"value":2426},"Requires a code release (maybe):",{"type":52,"value":2428}," unless you're doing all your frontend data collection with a tag manager updating your data collection will likely require a release of some kind with associated QA and processes surrounding release.",{"type":47,"tag":141,"props":2430,"children":2431},{},[2432,2437],{"type":47,"tag":88,"props":2433,"children":2434},{},[2435],{"type":52,"value":2436},"Harder to update data after collection:",{"type":52,"value":2438}," most platforms allow some degree of data management to help fix minor errors in the data. However, fixing the data is much harder after it is already collected and in downstream systems.",{"type":47,"tag":141,"props":2440,"children":2441},{},[2442,2447],{"type":47,"tag":88,"props":2443,"children":2444},{},[2445],{"type":52,"value":2446},"Typically no access to private data:",{"type":52,"value":2448}," frontend data collection is happening in the wild on client machines which typically means whatever data is available to pass into your data platform is available for anyone to snoop. For security reasons frontends only expose data that is safe to have publicly accessible which means any data collection on the frontend is limited to that information.",{"type":47,"tag":55,"props":2450,"children":2452},{"id":2451},"data-warehouse-ingestion",[2453],{"type":52,"value":2454},"Data Warehouse Ingestion",{"type":47,"tag":1057,"props":2456,"children":2458},{"id":2457},"what-are-the-benefits-1",[2459],{"type":52,"value":2366},{"type":47,"tag":137,"props":2461,"children":2462},{},[2463,2473],{"type":47,"tag":141,"props":2464,"children":2465},{},[2466,2471],{"type":47,"tag":88,"props":2467,"children":2468},{},[2469],{"type":52,"value":2470},"Post-hoc updates:",{"type":52,"value":2472}," because most of your data warehouse ingestion is going to be reliant on keys to join the data it also means you can change the data that you're joining onto those keys fairly easily.",{"type":47,"tag":141,"props":2474,"children":2475},{},[2476,2481],{"type":47,"tag":88,"props":2477,"children":2478},{},[2479],{"type":52,"value":2480},"Access to non-public data:",{"type":52,"value":2482}," one major benefit to data warehouse ingestion is that you're able to enrich your events with any non-public data you have. While the frontend is limited to just public data for security reasons your data warehouse can pull in any sensitive data because that data is not exposed to the public.",{"type":47,"tag":1057,"props":2484,"children":2486},{"id":2485},"what-are-the-drawbacks-1",[2487],{"type":52,"value":2415},{"type":47,"tag":137,"props":2489,"children":2490},{},[2491,2501,2511,2521],{"type":47,"tag":141,"props":2492,"children":2493},{},[2494,2499],{"type":47,"tag":88,"props":2495,"children":2496},{},[2497],{"type":52,"value":2498},"Data may not match what a user saw:",{"type":52,"value":2500}," Ingestion processes tend to be batches for efficiency but that also means the data in the ingested source has time to change and deviate from what a user actually experienced. This may mean the data isn't truly representative of a user's experience.",{"type":47,"tag":141,"props":2502,"children":2503},{},[2504,2509],{"type":47,"tag":88,"props":2505,"children":2506},{},[2507],{"type":52,"value":2508},"Centralized processing:",{"type":52,"value":2510}," the ingestion is all happening on your servers which means you will incur all of the compute cost and data transfer cost; doing too much ingestion can drive up your cloud computing bills.",{"type":47,"tag":141,"props":2512,"children":2513},{},[2514,2519],{"type":47,"tag":88,"props":2515,"children":2516},{},[2517],{"type":52,"value":2518},"Data only available in Data Warehouse:",{"type":52,"value":2520}," if you're using a visualization tool such as Amplitude, Adobe Analytics, Google Analytics, etc that tool won't have access to the full suite of information that is updated on the Data Warehouse side which means the reporting has to be done separately. Additionally, if you have a CDP or similar it means this enriched data cannot be used for creating audiences and is harder to use for triggering other CRM events.",{"type":47,"tag":141,"props":2522,"children":2523},{},[2524,2529],{"type":47,"tag":88,"props":2525,"children":2526},{},[2527],{"type":52,"value":2528},"Complexity:",{"type":52,"value":2530}," setting up the timed tasks and writing the appropriate ETL processes can be more complicated than the equivalent work on the frontend.",{"type":47,"tag":55,"props":2532,"children":2533},{"id":890},[2534],{"type":52,"value":893},{"type":47,"tag":48,"props":2536,"children":2537},{},[2538],{"type":52,"value":2539},"I hope this guide will give you a deeper understanding of your options for how to enrich your data. The next time this question comes up you can be more thoughtful in your planning. Whether you collect the data via Frontend Data Collection or Data Warehouse Ingestion, Anamap can help you map your data so analysts and stakeholders can more easily use the data to generate better insights.",{"title":8,"searchDepth":587,"depth":588,"links":2541},[2542,2543,2547,2551],{"id":2331,"depth":587,"text":2334},{"id":2357,"depth":587,"text":2360,"children":2544},[2545,2546],{"id":2363,"depth":588,"text":2366},{"id":2412,"depth":588,"text":2415},{"id":2451,"depth":587,"text":2454,"children":2548},[2549,2550],{"id":2457,"depth":588,"text":2366},{"id":2485,"depth":588,"text":2415},{"id":890,"depth":587,"text":893},"content:blog:frontend-vs-backend-data-collection.md","blog/frontend-vs-backend-data-collection.md","blog/frontend-vs-backend-data-collection",{"_path":2556,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":2557,"description":2558,"draft":7,"publicationDate":2559,"image":2560,"author":2561,"head":2562,"category":733,"body":2569,"_type":596,"_id":2764,"_source":598,"_file":2765,"_stem":2766,"_extension":601},"/blog/jumpstart-self-service-analytics","How You Can Jumpstart Self-Service Analytics","The core principles for getting the self-service analytics engine started are document, simplify, tool up, and train. Here are a few ways to get your business' stakeholders on the path to data driven decision-making.","2024-06-12","/images/blog/f052700f-c07e-4968-aab3-588fe584cdac.webp",{"id":14,"name":15,"role":16},{"meta":2563},[2564,2566,2567,2568],{"name":33,"content":2565},"analytics, self-service analytics, data driven, data democratization",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},{"type":44,"children":2570,"toc":2758},[2571,2575,2581,2586,2592,2597,2602,2608,2613,2618,2661,2673,2678,2690,2695,2700,2706],{"type":47,"tag":48,"props":2572,"children":2573},{},[2574],{"type":52,"value":2558},{"type":47,"tag":55,"props":2576,"children":2578},{"id":2577},"what-is-self-serve-analytics",[2579],{"type":52,"value":2580},"What is self-serve analytics?",{"type":47,"tag":48,"props":2582,"children":2583},{},[2584],{"type":52,"value":2585},"Self-service analytics (SSA) is the enablement of business stakeholders (non-analysts) to answer basic questions about the performance of the business themselves without having to rely on more technical peers. Fundamentally, self-service analytics is about streamlining your data collection and reporting so that it is easy to understand and utilize for people who aren't necessarily in the data day in and day out.",{"type":47,"tag":55,"props":2587,"children":2589},{"id":2588},"how-can-self-serve-analytics-help-my-business",[2590],{"type":52,"value":2591},"How can self-serve analytics help my business?",{"type":47,"tag":48,"props":2593,"children":2594},{},[2595],{"type":52,"value":2596},"It's easy to see the phrase \"self-service analytics'' and think \"that's just another buzzword\", however, nothing could be further from the truth. One of the tightest process bottlenecks in any organization is requesting work through the dedicated analytics team. Even for large enterprise organizations the analytics teams are traditionally too small to service every request for data the organization has. By making data easier to access and understand, stakeholders who would normally have to wait for their analysis to be completed can use the reporting and visualization tools themselves to get insights faster. Faster insights means faster decisions and benefits your organization's ability to respond and adapt to changing market conditions. Because these stakeholders now have the ability to get data for themselves when they need it rather than having the added friction of relying on an analyst for a simple request it means these stakeholders are more likely to have their decisions driven by data rather than guesses.",{"type":47,"tag":48,"props":2598,"children":2599},{},[2600],{"type":52,"value":2601},"Beyond making your business more data driven, democratizing your data has one other frequently overlooked benefit; it allows your analytics team to do less “low skill” reporting and spend more time focusing on “high skill” analysis like clustering and predictive models. This benefit by itself can be game changing for organizations because it can unlock new growth potential that was previously unavailable.",{"type":47,"tag":55,"props":2603,"children":2605},{"id":2604},"how-can-my-company-set-up-self-serve-analytics",[2606],{"type":52,"value":2607},"How can my company set up self-serve analytics?",{"type":47,"tag":48,"props":2609,"children":2610},{},[2611],{"type":52,"value":2612},"Given that analytics or business intelligence teams would have more time for interesting and brain stretching analysis you would think that many of them would be all aboard the self-service analytics train. The reality is that many analytics teams are so used to doing this low-level analysis that they worry they won't be useful if other stakeholders are able to do that reporting. They may couch that fear in statements like “we're worried about the integrity and consistency of company KPIs” or that stakeholders won't be able to do the reporting correctly. These sentiments are typically expressed because the analysts are worried about their job security. Given the typical response from many analytics teams, step one for rolling out self-serve analytics is to help the analysts in your organization see how it benefits them.",{"type":47,"tag":48,"props":2614,"children":2615},{},[2616],{"type":52,"value":2617},"Winning hearts and minds is a key part of the journey to self-serve analytics but it isn't the whole story. Making the data easy to use is the next big hurdle. Easy to use has a few different components but they roughly correspond to three tenants:",{"type":47,"tag":137,"props":2619,"children":2620},{},[2621,2631,2641,2651],{"type":47,"tag":141,"props":2622,"children":2623},{},[2624,2629],{"type":47,"tag":88,"props":2625,"children":2626},{},[2627],{"type":52,"value":2628},"Document:",{"type":52,"value":2630}," Provide good documentation about what events and attributes are available and what they are used for.",{"type":47,"tag":141,"props":2632,"children":2633},{},[2634,2639],{"type":47,"tag":88,"props":2635,"children":2636},{},[2637],{"type":52,"value":2638},"Simplify:",{"type":52,"value":2640}," Only expose the events and attributes that are absolutely necessary to limit confusion.",{"type":47,"tag":141,"props":2642,"children":2643},{},[2644,2649],{"type":47,"tag":88,"props":2645,"children":2646},{},[2647],{"type":52,"value":2648},"Tool Up:",{"type":52,"value":2650}," Pick tools that are easy to learn for visualizing data.",{"type":47,"tag":141,"props":2652,"children":2653},{},[2654,2659],{"type":47,"tag":88,"props":2655,"children":2656},{},[2657],{"type":52,"value":2658},"Train:",{"type":52,"value":2660}," Provide adequate ongoing training to stakeholders that covers 1, 2, and 3.",{"type":47,"tag":48,"props":2662,"children":2663},{},[2664,2666,2672],{"type":52,"value":2665},"Your documentation should include a full map of all your analytics events and attributes but there should be specific documentation for business stakeholders that only covers the essentials so they can focus on those; this is one area where Anamap can help. Anamap allows you to visually layout your customer experience so stakeholders understand the flow through your site or app and allows you to control which events and attributes are exposed to users. Having control over which attributes are shown to users can have a big impact on the adoption of SSA. If the documentation seems daunting and overly complicated users are likely to be overwhelmed or experience frustration finding important events resulting in disengagement from the self-serve analytics process. ",{"type":47,"tag":67,"props":2667,"children":2669},{"href":2668},"/",[2670],{"type":52,"value":2671},"Check out the free trial of Anamap to get started",{"type":52,"value":94},{"type":47,"tag":48,"props":2674,"children":2675},{},[2676],{"type":52,"value":2677},"As far as easy data visualization and reporting tools go I can recommend a few:\nAmplitude: it's expensive but I think it generally provides a good UI that is easy for non-analysts to use with a bit of training. My only gripe is that Amplitude needs a way to hide events and attributes for specific user groups so the dropdowns are more streamlined. Additionally, Amplitude can be pretty expensive.\nLooker: acquired by Google a while ago this is baked into the Google Analytics environment now. It provides a good interface for reporting but it can be a bit more complicated and similar to more advanced BI tools such as Tableau.\nAdobe Analytics: as much as I would love to move away from Adobe Analytics I have heard so many analysts and stakeholders say they like how the UI operates. The drag and drop nature and straight forward table interactions make this a potentially good choice. Again, the biggest problem here is cost. Adobe charges a pretty penny for a sizable analytics footprint; they also suffer from the inability to hide events and attributes for certain users.",{"type":47,"tag":48,"props":2679,"children":2680},{},[2681,2683,2688],{"type":52,"value":2682},"The last but arguably more important piece is training. Proper training is so often overlooked that I covered it in my ",{"type":47,"tag":67,"props":2684,"children":2685},{"href":1962},[2686],{"type":52,"value":2687},"common analytics mistakes",{"type":52,"value":2689}," post. For training the best option is to create small tailored training sessions for groups of business stakeholders within your organization and support those with a weekly drop-in Q&A or brown bag session. Most tools cannot be properly taught in a single session even with targeted training.",{"type":47,"tag":48,"props":2691,"children":2692},{},[2693],{"type":52,"value":2694},"For the initial course of training I would recommend at least three 2 hour long sessions all squeezed into a two week period if possible. Under ideal circumstances you would have short “homework” assignments for the users that could be walked through as part of the next training session. These training sessions and homeworks can be encouraged through the use of workplace rewards such as recognition, badges, etc.",{"type":47,"tag":48,"props":2696,"children":2697},{},[2698],{"type":52,"value":2699},"The weekly drop-in Q&A or brown bag session can provide a low barrier for users to come and get help if they forget specific aspects of a tool or your company data schema. This low barrier is important because it makes your self-serve analytics users more likely to come to you with problems and continue using the tool versus having a problem and just silently giving up.",{"type":47,"tag":55,"props":2701,"children":2703},{"id":2702},"what-are-some-tools-to-help-with-democratizing-my-data",[2704],{"type":52,"value":2705},"What are some tools to help with democratizing my data?",{"type":47,"tag":137,"props":2707,"children":2708},{},[2709,2732,2745],{"type":47,"tag":141,"props":2710,"children":2711},{},[2712,2714],{"type":52,"value":2713},"Documentation:\n",{"type":47,"tag":137,"props":2715,"children":2716},{},[2717,2722,2727],{"type":47,"tag":141,"props":2718,"children":2719},{},[2720],{"type":52,"value":2721},"Not surprisingly we recommend Anamap here because we built it for the express use of making it easier for everyone in your business to understand everything being tracked in a visual analytics map.",{"type":47,"tag":141,"props":2723,"children":2724},{},[2725],{"type":52,"value":2726},"If you're not sold on Anamap the next best option is a wiki using Confluence, SharePoint, Notion, or something similar. These wiki pages can have links to other pages which can be maintained by separate teams to distribute the work of maintenance.",{"type":47,"tag":141,"props":2728,"children":2729},{},[2730],{"type":52,"value":2731},"Your last option, though it's not recommended, is a spreadsheet. Having seen just about every format of analytics tracking spreadsheet I can tell you that if you use one you're going to be hard pressed to get self-serve analytics off the ground. Spreadsheets are hard to maintain as your data schema gets larger and more complicated and they are challenging to read even for technical users.",{"type":47,"tag":141,"props":2733,"children":2734},{},[2735,2737],{"type":52,"value":2736},"Easier to learn reporting tools:\n",{"type":47,"tag":137,"props":2738,"children":2739},{},[2740],{"type":47,"tag":141,"props":2741,"children":2742},{},[2743],{"type":52,"value":2744},"Amplitude, Looker, or Adobe Analytics would be my choices here. There are probably other great options out there that I haven’t used just yet.",{"type":47,"tag":141,"props":2746,"children":2747},{},[2748,2750],{"type":52,"value":2749},"Training incentive program:\n",{"type":47,"tag":137,"props":2751,"children":2752},{},[2753],{"type":47,"tag":141,"props":2754,"children":2755},{},[2756],{"type":52,"value":2757},"Vantage Circle, Work Human, Kudos, and Cool Leaf have an internal reward / recognition system that allows employees to earn points as well as badges. It can be used to help incentivize company employees to perform specific actions that further your self-serve analytics initiative.",{"title":8,"searchDepth":587,"depth":588,"links":2759},[2760,2761,2762,2763],{"id":2577,"depth":587,"text":2580},{"id":2588,"depth":587,"text":2591},{"id":2604,"depth":587,"text":2607},{"id":2702,"depth":587,"text":2705},"content:blog:jumpstart-self-service-analytics.md","blog/jumpstart-self-service-analytics.md","blog/jumpstart-self-service-analytics",{"_path":2768,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":2769,"description":2770,"draft":7,"publicationDate":2771,"image":2772,"author":2773,"head":2774,"category":2781,"body":2782,"_type":596,"_id":3194,"_source":598,"_file":3195,"_stem":3196,"_extension":601},"/blog/land-an-analyst-job","Land an Analyst Job with No Experience","Discover practical steps to break into an analyst career, from picking the right role and building a standout resume to finding job leads and nailing the interview—no prior experience needed.","2024-10-16","/images/blog/14abfe0b-414d-4edc-bfcd-76c2e803d108.webp",{"id":14,"name":15,"role":16},{"meta":2775},[2776,2778,2779,2780],{"name":33,"content":2777},"analytics, getting a job, hiring, resume, interviews, recruiting",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},"Recruiting",{"type":44,"children":2783,"toc":3186},[2784,2790,2795,2801,2806,2859,2880,2893,2926,2932,2937,3005,3011,3030,3069,3075,3080,3122,3128,3133,3175,3181],{"type":47,"tag":2785,"props":2786,"children":2788},"h1",{"id":2787},"land-an-analyst-job-with-no-experience",[2789],{"type":52,"value":2769},{"type":47,"tag":48,"props":2791,"children":2792},{},[2793],{"type":52,"value":2794},"Breaking into the world of analytics can feel like a daunting task, especially when you don’t have prior experience. But with the right approach, you can launch your career and land that first analyst job. Whether you’re just starting out or switching careers, this guide will help you navigate the path to becoming an analyst, from understanding the different types of roles to nailing the interview.",{"type":47,"tag":55,"props":2796,"children":2798},{"id":2797},"what-kinds-of-analyst-jobs-are-there",[2799],{"type":52,"value":2800},"What Kinds of Analyst Jobs Are There?",{"type":47,"tag":48,"props":2802,"children":2803},{},[2804],{"type":52,"value":2805},"The term \"analyst\" covers a lot of ground. Different companies look for analysts with varying specializations, and understanding the landscape can help you narrow your focus. Here are some of the most common types of analyst roles you’ll come across:",{"type":47,"tag":522,"props":2807,"children":2808},{},[2809,2819,2829,2839,2849],{"type":47,"tag":141,"props":2810,"children":2811},{},[2812,2817],{"type":47,"tag":88,"props":2813,"children":2814},{},[2815],{"type":52,"value":2816},"Data Analyst",{"type":52,"value":2818}," – Think of data analysts as the storytellers of numbers. They dig into datasets to uncover trends, draw insights, and support decision-making through data-driven storytelling.",{"type":47,"tag":141,"props":2820,"children":2821},{},[2822,2827],{"type":47,"tag":88,"props":2823,"children":2824},{},[2825],{"type":52,"value":2826},"Business Analyst",{"type":52,"value":2828}," – A business analyst bridges the gap between IT and business teams, focusing on process improvements and solutions that drive better business outcomes.",{"type":47,"tag":141,"props":2830,"children":2831},{},[2832,2837],{"type":47,"tag":88,"props":2833,"children":2834},{},[2835],{"type":52,"value":2836},"Financial Analyst",{"type":52,"value":2838}," – This role revolves around financial modeling, budgeting, and forecasting. Financial analysts are common in banks, investment firms, and large corporations.",{"type":47,"tag":141,"props":2840,"children":2841},{},[2842,2847],{"type":47,"tag":88,"props":2843,"children":2844},{},[2845],{"type":52,"value":2846},"Marketing Analyst",{"type":52,"value":2848}," – If digital marketing and customer behavior pique your interest, marketing analysts use data to fine-tune campaigns and maximize ROI.",{"type":47,"tag":141,"props":2850,"children":2851},{},[2852,2857],{"type":47,"tag":88,"props":2853,"children":2854},{},[2855],{"type":52,"value":2856},"Operations Analyst",{"type":52,"value":2858}," – An operations analyst dives into the nitty-gritty of business processes, optimizing operations, supply chains, and logistics.",{"type":47,"tag":48,"props":2860,"children":2861},{},[2862,2867,2871,2873,2878],{"type":47,"tag":88,"props":2863,"children":2864},{},[2865],{"type":52,"value":2866},"Choose Your Path",{"type":47,"tag":2868,"props":2869,"children":2870},"br",{},[],{"type":52,"value":2872},"\nStart by asking yourself: ",{"type":47,"tag":178,"props":2874,"children":2875},{},[2876],{"type":52,"value":2877},"Which type of analytics work excites me?",{"type":52,"value":2879}," The answer to this question will guide your next steps. Once you’ve identified your area of interest, research companies that hire for these roles and industries where they’re most prevalent. For instance, marketing analysts are in demand at digital agencies and e-commerce brands, while financial analysts find their fit in banks and investment firms.",{"type":47,"tag":48,"props":2881,"children":2882},{},[2883,2888,2891],{"type":47,"tag":88,"props":2884,"children":2885},{},[2886],{"type":52,"value":2887},"Entry-Level Roles and Where to Look",{"type":47,"tag":2868,"props":2889,"children":2890},{},[],{"type":52,"value":2892},"\nYour first job title may include words like \"junior,\" \"coordinator,\" or \"associate.\" Here are some common entry-level roles:",{"type":47,"tag":137,"props":2894,"children":2895},{},[2896,2906,2916],{"type":47,"tag":141,"props":2897,"children":2898},{},[2899,2904],{"type":47,"tag":88,"props":2900,"children":2901},{},[2902],{"type":52,"value":2903},"Junior Data Analyst",{"type":52,"value":2905}," – Frequently found in tech companies, marketing agencies, and the healthcare industry.",{"type":47,"tag":141,"props":2907,"children":2908},{},[2909,2914],{"type":47,"tag":88,"props":2910,"children":2911},{},[2912],{"type":52,"value":2913},"Operations Coordinator",{"type":52,"value":2915}," – Typically associated with logistics companies and retail businesses.",{"type":47,"tag":141,"props":2917,"children":2918},{},[2919,2924],{"type":47,"tag":88,"props":2920,"children":2921},{},[2922],{"type":52,"value":2923},"Marketing Data Assistant",{"type":52,"value":2925}," – E-commerce firms and larger marketing departments often seek candidates for this role.",{"type":47,"tag":55,"props":2927,"children":2929},{"id":2928},"what-do-i-put-on-my-resume",[2930],{"type":52,"value":2931},"What Do I Put on My Resume?",{"type":47,"tag":48,"props":2933,"children":2934},{},[2935],{"type":52,"value":2936},"Without direct experience, your resume needs to demonstrate relevant skills, coursework, and projects that show you’re ready for the job. Let’s break down the key elements:",{"type":47,"tag":522,"props":2938,"children":2939},{},[2940,2953,2966,2979,2992],{"type":47,"tag":141,"props":2941,"children":2942},{},[2943,2948,2951],{"type":47,"tag":88,"props":2944,"children":2945},{},[2946],{"type":52,"value":2947},"Start with Education",{"type":47,"tag":2868,"props":2949,"children":2950},{},[],{"type":52,"value":2952},"\nYour degree should be at the top, emphasizing relevant coursework and any notable projects. Focus on the tools and methods you used, like Excel for financial modeling, SQL for querying databases, or Tableau for creating data visualizations.",{"type":47,"tag":141,"props":2954,"children":2955},{},[2956,2961,2964],{"type":47,"tag":88,"props":2957,"children":2958},{},[2959],{"type":52,"value":2960},"Project Experience Speaks Volumes",{"type":47,"tag":2868,"props":2962,"children":2963},{},[],{"type":52,"value":2965},"\nWhether they’re class assignments or personal projects, showcase work that demonstrates your skills in action. Use concrete examples, such as analyzing customer data to find patterns or building a financial model to forecast growth.",{"type":47,"tag":141,"props":2967,"children":2968},{},[2969,2974,2977],{"type":47,"tag":88,"props":2970,"children":2971},{},[2972],{"type":52,"value":2973},"Boost Your Profile with Kaggle",{"type":47,"tag":2868,"props":2975,"children":2976},{},[],{"type":52,"value":2978},"\nKaggle competitions are an excellent way to gain hands-on experience. Look for challenges related to your desired role, whether that’s predictive modeling, marketing analytics, or business forecasting. Not only does this help you build skills, but it’s also a great addition to your resume.",{"type":47,"tag":141,"props":2980,"children":2981},{},[2982,2987,2990],{"type":47,"tag":88,"props":2983,"children":2984},{},[2985],{"type":52,"value":2986},"Highlight Internships (If Applicable)",{"type":47,"tag":2868,"props":2988,"children":2989},{},[],{"type":52,"value":2991},"\nIf you’ve completed an internship, describe the datasets you worked with and the outcomes of your analyses. Frame your work in terms of impact, such as \"Reduced customer churn by 10% through analysis of user behavior data.\"",{"type":47,"tag":141,"props":2993,"children":2994},{},[2995,3000,3003],{"type":47,"tag":88,"props":2996,"children":2997},{},[2998],{"type":52,"value":2999},"Add Relevant Experience, Even if It’s Not in Analytics",{"type":47,"tag":2868,"props":3001,"children":3002},{},[],{"type":52,"value":3004},"\nYour previous work may not have been in an analyst role, but chances are, you’ve gained valuable transferable skills. For example, retail jobs can showcase customer service skills, while roles in finance demonstrate attention to detail. Use your bullet points to highlight problem-solving abilities, teamwork, and industry-specific knowledge.",{"type":47,"tag":55,"props":3006,"children":3008},{"id":3007},"where-do-i-find-analyst-jobs",[3009],{"type":52,"value":3010},"Where do I Find Analyst Jobs?",{"type":47,"tag":48,"props":3012,"children":3013},{},[3014,3016,3021,3023,3028],{"type":52,"value":3015},"Finding the right job can be as much about ",{"type":47,"tag":178,"props":3017,"children":3018},{},[3019],{"type":52,"value":3020},"where",{"type":52,"value":3022}," you look as ",{"type":47,"tag":178,"props":3024,"children":3025},{},[3026],{"type":52,"value":3027},"how",{"type":52,"value":3029}," you look. Here’s how to maximize your chances:",{"type":47,"tag":522,"props":3031,"children":3032},{},[3033,3046,3056],{"type":47,"tag":141,"props":3034,"children":3035},{},[3036,3041,3044],{"type":47,"tag":88,"props":3037,"children":3038},{},[3039],{"type":52,"value":3040},"Start Networking on LinkedIn Early",{"type":47,"tag":2868,"props":3042,"children":3043},{},[],{"type":52,"value":3045},"\nDon’t wait until graduation to start building your professional network. Aim to grow your connections at least a year before you’re ready to apply for jobs. Reach out to people working in roles you aspire to and express your interest in their career path. Erin McGough’s advice from AdviceWithErin offers great templates for sending these kinds of messages.",{"type":47,"tag":141,"props":3047,"children":3048},{},[3049,3054],{"type":47,"tag":88,"props":3050,"children":3051},{},[3052],{"type":52,"value":3053},"Search Job Boards like LinkedIn, Glassdoor, Indeed, etc",{"type":52,"value":3055},"\nMost companies use job boards for posting their jobs these days. You should be able to find a wide variety of roles that are posted on any of these just using the keywords \"analyst\", \"analytics\", or \"data\". It's worth mentioning that that networking in tip number 1 from this section can also lead to jobs found directly from your connections on LinkedIn sharing them. If you target these you will have an above average chance of at least getting a first round interview.",{"type":47,"tag":141,"props":3057,"children":3058},{},[3059,3064,3067],{"type":47,"tag":88,"props":3060,"children":3061},{},[3062],{"type":52,"value":3063},"Use Glassdoor for Company Insights",{"type":47,"tag":2868,"props":3065,"children":3066},{},[],{"type":52,"value":3068},"\nGlassdoor isn’t just for finding job postings; it’s also a window into company culture. Look up companies you’re interested in, read reviews, and check out the salary ranges for analyst positions.",{"type":47,"tag":55,"props":3070,"children":3072},{"id":3071},"how-do-i-apply-for-analyst-roles",[3073],{"type":52,"value":3074},"How Do I Apply for Analyst Roles?",{"type":47,"tag":48,"props":3076,"children":3077},{},[3078],{"type":52,"value":3079},"When it’s time to start sending out applications, keep these tips in mind to improve your chances:",{"type":47,"tag":522,"props":3081,"children":3082},{},[3083,3096,3109],{"type":47,"tag":141,"props":3084,"children":3085},{},[3086,3091,3094],{"type":47,"tag":88,"props":3087,"children":3088},{},[3089],{"type":52,"value":3090},"Use Job Boards but Confirm on the Company’s Careers Page",{"type":47,"tag":2868,"props":3092,"children":3093},{},[],{"type":52,"value":3095},"\nJob boards like LinkedIn, Indeed, and Glassdoor are great for finding openings, but always confirm the job listing on the company’s official careers site. Sometimes third-party sites have outdated information, while the company’s site is likely to be up to date.",{"type":47,"tag":141,"props":3097,"children":3098},{},[3099,3104,3107],{"type":47,"tag":88,"props":3100,"children":3101},{},[3102],{"type":52,"value":3103},"Avoid Stale Job Postings",{"type":47,"tag":2868,"props":3105,"children":3106},{},[],{"type":52,"value":3108},"\nApply to the freshest listings you can find—ideally posted within the last two weeks. Older postings may mean that the company is already deep into the interview process. The exception here is for listings specifically for recent graduates, which companies may keep open for extended periods.",{"type":47,"tag":141,"props":3110,"children":3111},{},[3112,3117,3120],{"type":47,"tag":88,"props":3113,"children":3114},{},[3115],{"type":52,"value":3116},"Get a Referral If You Can",{"type":47,"tag":2868,"props":3118,"children":3119},{},[],{"type":52,"value":3121},"\nReferrals can significantly boost your application’s visibility. Reach out to someone who works at the company and, if possible, ask about their experience before requesting a referral. While referrals don’t guarantee an interview, they do get you a closer look.",{"type":47,"tag":55,"props":3123,"children":3125},{"id":3124},"how-do-i-prepare-for-an-analyst-job-interview",[3126],{"type":52,"value":3127},"How do I Prepare for an Analyst Job Interview?",{"type":47,"tag":48,"props":3129,"children":3130},{},[3131],{"type":52,"value":3132},"Landing an interview is a significant milestone. Now’s the time to prepare by following these steps:",{"type":47,"tag":522,"props":3134,"children":3135},{},[3136,3149,3162],{"type":47,"tag":141,"props":3137,"children":3138},{},[3139,3144,3147],{"type":47,"tag":88,"props":3140,"children":3141},{},[3142],{"type":52,"value":3143},"Leverage Glassdoor for Interview Questions",{"type":47,"tag":2868,"props":3145,"children":3146},{},[],{"type":52,"value":3148},"\nMany candidates share common interview questions on Glassdoor. Look up the company or role you’re interviewing for and practice your answers. Use the STAR method (Situation, Task, Action, Result) to structure your responses.",{"type":47,"tag":141,"props":3150,"children":3151},{},[3152,3157,3160],{"type":47,"tag":88,"props":3153,"children":3154},{},[3155],{"type":52,"value":3156},"Learn About the Company’s Revenue Streams",{"type":47,"tag":2868,"props":3158,"children":3159},{},[],{"type":52,"value":3161},"\nUnderstanding how the company makes money will show that you’ve done your homework. Are there multiple revenue streams? Which ones are the most profitable? This research can help you tailor your answers and ask informed questions during the interview.",{"type":47,"tag":141,"props":3163,"children":3164},{},[3165,3170,3173],{"type":47,"tag":88,"props":3166,"children":3167},{},[3168],{"type":52,"value":3169},"Know the Salary Range for the Role",{"type":47,"tag":2868,"props":3171,"children":3172},{},[],{"type":52,"value":3174},"\nResearching salary ranges for analyst roles can prepare you for compensation discussions. Getting a fair starting salary has a ripple effect on your future earnings, so don’t skip this step. Studies show that negotiating your first salary can impact your lifetime earnings by up to $1 million.",{"type":47,"tag":55,"props":3176,"children":3178},{"id":3177},"final-thoughts",[3179],{"type":52,"value":3180},"Final Thoughts",{"type":47,"tag":48,"props":3182,"children":3183},{},[3184],{"type":52,"value":3185},"The journey to landing your first analyst job might be challenging, but it’s also rewarding. Take the time to research your options, tailor your resume, network strategically, and prepare thoroughly for interviews. With persistence and a plan, you’ll find yourself well on the way to your first role in analytics.",{"title":8,"searchDepth":587,"depth":588,"links":3187},[3188,3189,3190,3191,3192,3193],{"id":2797,"depth":587,"text":2800},{"id":2928,"depth":587,"text":2931},{"id":3007,"depth":587,"text":3010},{"id":3071,"depth":587,"text":3074},{"id":3124,"depth":587,"text":3127},{"id":3177,"depth":587,"text":3180},"content:blog:land-an-analyst-job.md","blog/land-an-analyst-job.md","blog/land-an-analyst-job",{"_path":1313,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":3198,"description":3199,"draft":7,"publicationDate":3200,"updatedAt":3200,"image":3201,"author":3202,"ogTitle":3203,"ogDescription":3204,"twitterTitle":3205,"twitterDescription":3206,"keywords":3207,"tags":3208,"head":3214,"category":42,"body":3221,"_type":596,"_id":5014,"_source":598,"_file":5015,"_stem":5016,"_extension":601},"I Benchmarked 10 LLMs on Broken Analytics Data. Only 30% Delivered Value.","All 10 leading AI models achieved perfect API syntax. But when given intentionally broken GA4 data with 100% attribution failure, their analytical judgment varied dramatically. Here's the complete breakdown.","2026-01-28","/images/blog/llm-benchmark-hero.svg",{"id":14,"name":15,"role":16,"twitter":17},"LLM Analytics Benchmark: Which AI Models Can Handle Broken Data?","I tested Claude, GPT-5, Gemini, Grok, and DeepSeek on a GA4 property with 100% broken attribution. Only 3 of 10 models delivered actionable insights. See the full results.","🔬 I Benchmarked 10 LLMs on Broken Analytics Data","All 10 got perfect API syntax. But only 30% could actually help when the data was broken. Claude Opus won, but Grok at $0.03 was the best value.","LLM benchmark, AI analytics comparison, Claude vs GPT-5 vs Gemini, best AI for Google Analytics, GA4 AI assistant, analytics AI hallucination, LLM data quality, Claude Opus 4.5 review, GPT-5 analytics test, Gemini Flash comparison, Grok AI benchmark, DeepSeek analytics, AI marketing attribution, machine learning analytics",[3209,25,930,931,932,3210,3211,928,3212,3213],"LLM benchmark","Grok","DeepSeek","data quality","model comparison",{"meta":3215},[3216,3218,3219,3220],{"name":33,"content":3217},"LLM benchmark, AI analytics comparison, Claude vs GPT-5 vs Gemini, best AI for Google Analytics, GA4 AI assistant, analytics AI hallucination, LLM data quality, Claude Opus 4.5 review, GPT-5 analytics test, Gemini Flash comparison, Grok AI benchmark, DeepSeek analytics",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},{"type":44,"children":3222,"toc":4976},[3223,3229,3286,3298,3315,3325,3331,3335,3343,3349,3354,3805,3822,3828,3834,3839,3846,3863,3868,3879,3887,3915,3925,3931,3946,3951,3962,3970,3993,3999,4014,4019,4030,4038,4061,4067,4072,4077,4092,4097,4108,4113,4118,4133,4138,4149,4154,4159,4173,4178,4189,4194,4199,4214,4219,4230,4235,4241,4246,4251,4265,4270,4281,4291,4296,4301,4316,4321,4332,4341,4346,4360,4365,4376,4381,4387,4393,4508,4513,4536,4542,4547,4590,4595,4601,4607,4615,4638,4646,4674,4682,4710,4713,4730,4734,4740,4757,4763,4768,4774,4779,4812,4818,4823,4855,4861,4866,4894,4900,4905,4909,4912,4916],{"type":47,"tag":55,"props":3224,"children":3226},{"id":3225},"the-setup-a-test-most-models-would-fail",[3227],{"type":52,"value":3228},"The Setup: A Test Most Models Would Fail",{"type":47,"tag":965,"props":3230,"children":3232},{"icon":967,"title":3231,"type":890},"Key Takeaways",[3233],{"type":47,"tag":137,"props":3234,"children":3235},{},[3236,3246,3256,3266,3276],{"type":47,"tag":141,"props":3237,"children":3238},{},[3239,3244],{"type":47,"tag":88,"props":3240,"children":3241},{},[3242],{"type":52,"value":3243},"All 10 models",{"type":52,"value":3245}," achieved perfect API syntax and valid GA4 queries",{"type":47,"tag":141,"props":3247,"children":3248},{},[3249,3254],{"type":47,"tag":88,"props":3250,"children":3251},{},[3252],{"type":52,"value":3253},"Only 30%",{"type":52,"value":3255}," provided actionable insights when data was broken",{"type":47,"tag":141,"props":3257,"children":3258},{},[3259,3264],{"type":47,"tag":88,"props":3260,"children":3261},{},[3262],{"type":52,"value":3263},"Best overall:",{"type":52,"value":3265}," Claude Opus 4.5 ($1.30) found workarounds and extracted real value",{"type":47,"tag":141,"props":3267,"children":3268},{},[3269,3274],{"type":47,"tag":88,"props":3270,"children":3271},{},[3272],{"type":52,"value":3273},"Best value:",{"type":52,"value":3275}," Grok 4.1 Fast ($0.03) delivered solid analysis at 1/40th the cost",{"type":47,"tag":141,"props":3277,"children":3278},{},[3279,3284],{"type":47,"tag":88,"props":3280,"children":3281},{},[3282],{"type":52,"value":3283},"Dangerous:",{"type":52,"value":3285}," 30% hallucinated or gave misleading recommendations",{"type":47,"tag":48,"props":3287,"children":3288},{},[3289,3291,3296],{"type":52,"value":3290},"I benchmarked 10 leading LLMs on a deceptively simple analytics question using the Anamap AI library. But the data had a critical flaw: ",{"type":47,"tag":88,"props":3292,"children":3293},{},[3294],{"type":52,"value":3295},"100% of traffic showed as \"(not set)\"",{"type":52,"value":3297}," with zero conversion attribution.",{"type":47,"tag":48,"props":3299,"children":3300},{},[3301,3303,3308,3310],{"type":52,"value":3302},"This is a scenario every analytics professional dreads, and one that ",{"type":47,"tag":67,"props":3304,"children":3305},{"href":1962},[3306],{"type":52,"value":3307},"happens more often than you'd think",{"type":52,"value":3309},". The question I asked: ",{"type":47,"tag":178,"props":3311,"children":3312},{},[3313],{"type":52,"value":3314},"\"Which traffic sources and landing pages are driving our highest-value users, and where should we double down our marketing investment?\"",{"type":47,"tag":48,"props":3316,"children":3317},{},[3318,3320],{"type":52,"value":3319},"Every model got the technical part right: perfect API syntax, valid GA4 field names, clean query structures. But technical accuracy is table stakes. The real test was: ",{"type":47,"tag":88,"props":3321,"children":3322},{},[3323],{"type":52,"value":3324},"Would they catch the broken data before giving recommendations?",{"type":47,"tag":55,"props":3326,"children":3328},{"id":3327},"the-results-technical-accuracy-analytical-value",[3329],{"type":52,"value":3330},"The Results: Technical Accuracy ≠ Analytical Value",{"type":47,"tag":3332,"props":3333,"children":3334},"llm-benchmark-visualization",{},[],{"type":47,"tag":1245,"props":3336,"children":3342},{"heading":3337,"icon":3338,"primary-link":1249,"primary-text":1250,"secondary-link":3339,"secondary-text":3340,"subheading":3341},"Want AI analytics that handles broken data gracefully?","mdi-chart-box-outline","#newsletter","Subscribe for Updates","See how Anamap detects data quality issues before they mislead your team.",[],{"type":47,"tag":55,"props":3344,"children":3346},{"id":3345},"the-full-leaderboard",[3347],{"type":52,"value":3348},"The Full Leaderboard",{"type":47,"tag":48,"props":3350,"children":3351},{},[3352],{"type":52,"value":3353},"Here's how each model performed across all metrics:",{"type":47,"tag":1363,"props":3355,"children":3356},{},[3357,3397],{"type":47,"tag":1367,"props":3358,"children":3359},{},[3360],{"type":47,"tag":1371,"props":3361,"children":3362},{},[3363,3369,3374,3378,3383,3387,3392],{"type":47,"tag":1375,"props":3364,"children":3366},{"align":3365},"center",[3367],{"type":52,"value":3368},"Rank",{"type":47,"tag":1375,"props":3370,"children":3371},{},[3372],{"type":52,"value":3373},"Model",{"type":47,"tag":1375,"props":3375,"children":3376},{},[3377],{"type":52,"value":1629},{"type":47,"tag":1375,"props":3379,"children":3380},{},[3381],{"type":52,"value":3382},"Time",{"type":47,"tag":1375,"props":3384,"children":3385},{},[3386],{"type":52,"value":1485},{"type":47,"tag":1375,"props":3388,"children":3389},{},[3390],{"type":52,"value":3391},"Tokens",{"type":47,"tag":1375,"props":3393,"children":3394},{},[3395],{"type":52,"value":3396},"Data Quality Handling",{"type":47,"tag":1391,"props":3398,"children":3399},{},[3400,3442,3483,3523,3563,3604,3644,3684,3725,3765],{"type":47,"tag":1371,"props":3401,"children":3402},{},[3403,3408,3417,3422,3427,3432,3437],{"type":47,"tag":1398,"props":3404,"children":3405},{"align":3365},[3406],{"type":52,"value":3407},"1",{"type":47,"tag":1398,"props":3409,"children":3410},{},[3411],{"type":47,"tag":67,"props":3412,"children":3414},{"href":3413},"#1-claude-opus-45",[3415],{"type":52,"value":3416},"Claude Opus 4.5",{"type":47,"tag":1398,"props":3418,"children":3419},{},[3420],{"type":52,"value":3421},"🏆 excellent",{"type":47,"tag":1398,"props":3423,"children":3424},{},[3425],{"type":52,"value":3426},"96s",{"type":47,"tag":1398,"props":3428,"children":3429},{},[3430],{"type":52,"value":3431},"$1.30",{"type":47,"tag":1398,"props":3433,"children":3434},{},[3435],{"type":52,"value":3436},"240K",{"type":47,"tag":1398,"props":3438,"children":3439},{},[3440],{"type":52,"value":3441},"✅✅ Best analysis",{"type":47,"tag":1371,"props":3443,"children":3444},{},[3445,3450,3459,3463,3468,3473,3478],{"type":47,"tag":1398,"props":3446,"children":3447},{"align":3365},[3448],{"type":52,"value":3449},"2",{"type":47,"tag":1398,"props":3451,"children":3452},{},[3453],{"type":47,"tag":67,"props":3454,"children":3456},{"href":3455},"#2-claude-sonnet-45",[3457],{"type":52,"value":3458},"Claude Sonnet 4.5",{"type":47,"tag":1398,"props":3460,"children":3461},{},[3462],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3464,"children":3465},{},[3466],{"type":52,"value":3467},"124s",{"type":47,"tag":1398,"props":3469,"children":3470},{},[3471],{"type":52,"value":3472},"$0.66",{"type":47,"tag":1398,"props":3474,"children":3475},{},[3476],{"type":52,"value":3477},"191K",{"type":47,"tag":1398,"props":3479,"children":3480},{},[3481],{"type":52,"value":3482},"✅✅ Clear pivot",{"type":47,"tag":1371,"props":3484,"children":3485},{},[3486,3491,3499,3503,3508,3513,3518],{"type":47,"tag":1398,"props":3487,"children":3488},{"align":3365},[3489],{"type":52,"value":3490},"3",{"type":47,"tag":1398,"props":3492,"children":3493},{},[3494],{"type":47,"tag":67,"props":3495,"children":3497},{"href":3496},"#3-grok-41-fast",[3498],{"type":52,"value":1185},{"type":47,"tag":1398,"props":3500,"children":3501},{},[3502],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3504,"children":3505},{},[3506],{"type":52,"value":3507},"83s",{"type":47,"tag":1398,"props":3509,"children":3510},{},[3511],{"type":52,"value":3512},"$0.03",{"type":47,"tag":1398,"props":3514,"children":3515},{},[3516],{"type":52,"value":3517},"122K",{"type":47,"tag":1398,"props":3519,"children":3520},{},[3521],{"type":52,"value":3522},"✅ Found signal",{"type":47,"tag":1371,"props":3524,"children":3525},{},[3526,3531,3539,3543,3548,3553,3558],{"type":47,"tag":1398,"props":3527,"children":3528},{"align":3365},[3529],{"type":52,"value":3530},"4",{"type":47,"tag":1398,"props":3532,"children":3533},{},[3534],{"type":47,"tag":67,"props":3535,"children":3537},{"href":3536},"#gpt-5",[3538],{"type":52,"value":931},{"type":47,"tag":1398,"props":3540,"children":3541},{},[3542],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3544,"children":3545},{},[3546],{"type":52,"value":3547},"163s",{"type":47,"tag":1398,"props":3549,"children":3550},{},[3551],{"type":52,"value":3552},"$0.24",{"type":47,"tag":1398,"props":3554,"children":3555},{},[3556],{"type":52,"value":3557},"123K",{"type":47,"tag":1398,"props":3559,"children":3560},{},[3561],{"type":52,"value":3562},"✅ Diagnostic only",{"type":47,"tag":1371,"props":3564,"children":3565},{},[3566,3571,3580,3584,3589,3594,3599],{"type":47,"tag":1398,"props":3567,"children":3568},{"align":3365},[3569],{"type":52,"value":3570},"5",{"type":47,"tag":1398,"props":3572,"children":3573},{},[3574],{"type":47,"tag":67,"props":3575,"children":3577},{"href":3576},"#gemini-25-flash",[3578],{"type":52,"value":3579},"Gemini 2.5 Flash",{"type":47,"tag":1398,"props":3581,"children":3582},{},[3583],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3585,"children":3586},{},[3587],{"type":52,"value":3588},"27s",{"type":47,"tag":1398,"props":3590,"children":3591},{},[3592],{"type":52,"value":3593},"$0.15",{"type":47,"tag":1398,"props":3595,"children":3596},{},[3597],{"type":52,"value":3598},"449K",{"type":47,"tag":1398,"props":3600,"children":3601},{},[3602],{"type":52,"value":3603},"✅ Identified issue",{"type":47,"tag":1371,"props":3605,"children":3606},{},[3607,3612,3621,3625,3630,3634,3639],{"type":47,"tag":1398,"props":3608,"children":3609},{"align":3365},[3610],{"type":52,"value":3611},"6",{"type":47,"tag":1398,"props":3613,"children":3614},{},[3615],{"type":47,"tag":67,"props":3616,"children":3618},{"href":3617},"#deepseek-v32",[3619],{"type":52,"value":3620},"DeepSeek V3.2",{"type":47,"tag":1398,"props":3622,"children":3623},{},[3624],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3626,"children":3627},{},[3628],{"type":52,"value":3629},"199s",{"type":47,"tag":1398,"props":3631,"children":3632},{},[3633],{"type":52,"value":3512},{"type":47,"tag":1398,"props":3635,"children":3636},{},[3637],{"type":52,"value":3638},"125K",{"type":47,"tag":1398,"props":3640,"children":3641},{},[3642],{"type":52,"value":3643},"✅ No next steps",{"type":47,"tag":1371,"props":3645,"children":3646},{},[3647,3652,3661,3665,3670,3675,3680],{"type":47,"tag":1398,"props":3648,"children":3649},{"align":3365},[3650],{"type":52,"value":3651},"7",{"type":47,"tag":1398,"props":3653,"children":3654},{},[3655],{"type":47,"tag":67,"props":3656,"children":3658},{"href":3657},"#grok-code-fast-1",[3659],{"type":52,"value":3660},"Grok Code Fast 1",{"type":47,"tag":1398,"props":3662,"children":3663},{},[3664],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3666,"children":3667},{},[3668],{"type":52,"value":3669},"28s",{"type":47,"tag":1398,"props":3671,"children":3672},{},[3673],{"type":52,"value":3674},"$0.02",{"type":47,"tag":1398,"props":3676,"children":3677},{},[3678],{"type":52,"value":3679},"78K",{"type":47,"tag":1398,"props":3681,"children":3682},{},[3683],{"type":52,"value":3603},{"type":47,"tag":1371,"props":3685,"children":3686},{},[3687,3692,3701,3705,3710,3715,3720],{"type":47,"tag":1398,"props":3688,"children":3689},{"align":3365},[3690],{"type":52,"value":3691},"8",{"type":47,"tag":1398,"props":3693,"children":3694},{},[3695],{"type":47,"tag":67,"props":3696,"children":3698},{"href":3697},"#gemini-3-flash-preview",[3699],{"type":52,"value":3700},"Gemini 3 Flash Preview",{"type":47,"tag":1398,"props":3702,"children":3703},{},[3704],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3706,"children":3707},{},[3708],{"type":52,"value":3709},"11s",{"type":47,"tag":1398,"props":3711,"children":3712},{},[3713],{"type":52,"value":3714},"$0.05",{"type":47,"tag":1398,"props":3716,"children":3717},{},[3718],{"type":52,"value":3719},"87K",{"type":47,"tag":1398,"props":3721,"children":3722},{},[3723],{"type":52,"value":3724},"⚠️ Weak caveat",{"type":47,"tag":1371,"props":3726,"children":3727},{},[3728,3733,3742,3746,3751,3755,3760],{"type":47,"tag":1398,"props":3729,"children":3730},{"align":3365},[3731],{"type":52,"value":3732},"9",{"type":47,"tag":1398,"props":3734,"children":3735},{},[3736],{"type":47,"tag":67,"props":3737,"children":3739},{"href":3738},"#gpt-5-mini",[3740],{"type":52,"value":3741},"GPT-5 Mini",{"type":47,"tag":1398,"props":3743,"children":3744},{},[3745],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3747,"children":3748},{},[3749],{"type":52,"value":3750},"141s",{"type":47,"tag":1398,"props":3752,"children":3753},{},[3754],{"type":52,"value":3714},{"type":47,"tag":1398,"props":3756,"children":3757},{},[3758],{"type":52,"value":3759},"127K",{"type":47,"tag":1398,"props":3761,"children":3762},{},[3763],{"type":52,"value":3764},"❌ Misleading",{"type":47,"tag":1371,"props":3766,"children":3767},{},[3768,3773,3782,3786,3791,3795,3800],{"type":47,"tag":1398,"props":3769,"children":3770},{"align":3365},[3771],{"type":52,"value":3772},"10",{"type":47,"tag":1398,"props":3774,"children":3775},{},[3776],{"type":47,"tag":67,"props":3777,"children":3779},{"href":3778},"#gemini-25-flash-lite",[3780],{"type":52,"value":3781},"Gemini 2.5 Flash Lite",{"type":47,"tag":1398,"props":3783,"children":3784},{},[3785],{"type":52,"value":3421},{"type":47,"tag":1398,"props":3787,"children":3788},{},[3789],{"type":52,"value":3790},"48s",{"type":47,"tag":1398,"props":3792,"children":3793},{},[3794],{"type":52,"value":3674},{"type":47,"tag":1398,"props":3796,"children":3797},{},[3798],{"type":52,"value":3799},"170K",{"type":47,"tag":1398,"props":3801,"children":3802},{},[3803],{"type":52,"value":3804},"❌ Hallucinated",{"type":47,"tag":48,"props":3806,"children":3807},{},[3808,3813,3815,3820],{"type":47,"tag":88,"props":3809,"children":3810},{},[3811],{"type":52,"value":3812},"Note:",{"type":52,"value":3814}," All models achieved \"excellent\" quality ratings for ",{"type":47,"tag":88,"props":3816,"children":3817},{},[3818],{"type":52,"value":3819},"technical execution",{"type":52,"value":3821}," (valid GA4 queries, proper API syntax). The differentiation came entirely from how they handled the data quality problem.",{"type":47,"tag":55,"props":3823,"children":3825},{"id":3824},"category-breakdown",[3826],{"type":52,"value":3827},"Category Breakdown",{"type":47,"tag":1057,"props":3829,"children":3831},{"id":3830},"models-that-delivered-value-30",[3832],{"type":52,"value":3833},"🏆 Models That Delivered Value (30%)",{"type":47,"tag":48,"props":3835,"children":3836},{},[3837],{"type":52,"value":3838},"These models didn't just identify the problem. They provided actionable workarounds and next steps.",{"type":47,"tag":3840,"props":3841,"children":3843},"h4",{"id":3842},"_1-claude-opus-45",[3844],{"type":52,"value":3845},"1. Claude Opus 4.5",{"type":47,"tag":48,"props":3847,"children":3848},{},[3849,3854,3856,3861],{"type":47,"tag":88,"props":3850,"children":3851},{},[3852],{"type":52,"value":3853},"Cost:",{"type":52,"value":3855}," $1.30 | ",{"type":47,"tag":88,"props":3857,"children":3858},{},[3859],{"type":52,"value":3860},"Time:",{"type":52,"value":3862}," 96s",{"type":47,"tag":48,"props":3864,"children":3865},{},[3866],{"type":52,"value":3867},"Claude Opus didn't just flag the attribution failure. It pivoted to extract genuine insights from available data:",{"type":47,"tag":755,"props":3869,"children":3870},{},[3871],{"type":47,"tag":48,"props":3872,"children":3873},{},[3874],{"type":47,"tag":178,"props":3875,"children":3876},{},[3877],{"type":52,"value":3878},"\"Google organic search and LinkedIn are your two highest-value external traffic sources, driving engaged users to key conversion pages.\"",{"type":47,"tag":48,"props":3880,"children":3881},{},[3882],{"type":47,"tag":88,"props":3883,"children":3884},{},[3885],{"type":52,"value":3886},"What made it exceptional:",{"type":47,"tag":137,"props":3888,"children":3889},{},[3890,3895,3900,3905,3910],{"type":47,"tag":141,"props":3891,"children":3892},{},[3893],{"type":52,"value":3894},"Called out 100% \"(not set)\" attribution immediately",{"type":47,"tag":141,"props":3896,"children":3897},{},[3898],{"type":52,"value":3899},"Pivoted to pageReferrer analysis as a workaround",{"type":47,"tag":141,"props":3901,"children":3902},{},[3903],{"type":52,"value":3904},"Found signal in landing page performance and funnel paths",{"type":47,"tag":141,"props":3906,"children":3907},{},[3908],{"type":52,"value":3909},"Provided specific, actionable fixes for the tracking implementation",{"type":47,"tag":141,"props":3911,"children":3912},{},[3913],{"type":52,"value":3914},"Generated 6 detailed charts including traffic flow analysis",{"type":47,"tag":48,"props":3916,"children":3917},{},[3918,3923],{"type":47,"tag":88,"props":3919,"children":3920},{},[3921],{"type":52,"value":3922},"Key insight discovered:",{"type":52,"value":3924}," The /features page had 99.2% engagement rate and served as the primary gateway to pricing, a critical finding despite broken source attribution.",{"type":47,"tag":3840,"props":3926,"children":3928},{"id":3927},"_2-claude-sonnet-45",[3929],{"type":52,"value":3930},"2. Claude Sonnet 4.5",{"type":47,"tag":48,"props":3932,"children":3933},{},[3934,3938,3940,3944],{"type":47,"tag":88,"props":3935,"children":3936},{},[3937],{"type":52,"value":3853},{"type":52,"value":3939}," $0.66 | ",{"type":47,"tag":88,"props":3941,"children":3942},{},[3943],{"type":52,"value":3860},{"type":52,"value":3945}," 124s",{"type":47,"tag":48,"props":3947,"children":3948},{},[3949],{"type":52,"value":3950},"Sonnet led with the data quality issue but immediately pivoted to what IS working:",{"type":47,"tag":755,"props":3952,"children":3953},{},[3954],{"type":47,"tag":48,"props":3955,"children":3956},{},[3957],{"type":47,"tag":178,"props":3958,"children":3959},{},[3960],{"type":52,"value":3961},"\"Critical data quality issue detected: 100% of traffic shows '(not set)' for source/medium attribution... However, landing page data reveals strong engagement patterns.\"",{"type":47,"tag":48,"props":3963,"children":3964},{},[3965],{"type":47,"tag":88,"props":3966,"children":3967},{},[3968],{"type":52,"value":3969},"Standout analysis:",{"type":47,"tag":137,"props":3971,"children":3972},{},[3973,3978,3983,3988],{"type":47,"tag":141,"props":3974,"children":3975},{},[3976],{"type":52,"value":3977},"Identified 73% signup funnel abandonment rate (signup_started → sign_up)",{"type":47,"tag":141,"props":3979,"children":3980},{},[3981],{"type":52,"value":3982},"Found 414K subscription upgrades that couldn't be attributed",{"type":47,"tag":141,"props":3984,"children":3985},{},[3986],{"type":52,"value":3987},"Provided concrete implementation fixes",{"type":47,"tag":141,"props":3989,"children":3990},{},[3991],{"type":52,"value":3992},"Suggested proxy approaches while tracking is fixed",{"type":47,"tag":3840,"props":3994,"children":3996},{"id":3995},"_3-grok-41-fast",[3997],{"type":52,"value":3998},"3. Grok 4.1 Fast",{"type":47,"tag":48,"props":4000,"children":4001},{},[4002,4006,4008,4012],{"type":47,"tag":88,"props":4003,"children":4004},{},[4005],{"type":52,"value":3853},{"type":52,"value":4007}," $0.03 | ",{"type":47,"tag":88,"props":4009,"children":4010},{},[4011],{"type":52,"value":3860},{"type":52,"value":4013}," 83s",{"type":47,"tag":48,"props":4015,"children":4016},{},[4017],{"type":52,"value":4018},"The best value player. Despite the lowest cost among actionable models, Grok found real signal:",{"type":47,"tag":755,"props":4020,"children":4021},{},[4022],{"type":47,"tag":48,"props":4023,"children":4024},{},[4025],{"type":47,"tag":178,"props":4026,"children":4027},{},[4028],{"type":52,"value":4029},"\"Traffic sources are untrackable (100% '(not set)'), masking external marketing ROI... Subscription upgrades occur almost exclusively from in-app landing pages like /dashboard (28%) and /features (8%), signaling strong product-led growth.\"",{"type":47,"tag":48,"props":4031,"children":4032},{},[4033],{"type":47,"tag":88,"props":4034,"children":4035},{},[4036],{"type":52,"value":4037},"What it found:",{"type":47,"tag":137,"props":4039,"children":4040},{},[4041,4046,4051,4056],{"type":47,"tag":141,"props":4042,"children":4043},{},[4044],{"type":52,"value":4045},"Recognized product-led growth patterns from in-app upgrade paths",{"type":47,"tag":141,"props":4047,"children":4048},{},[4049],{"type":52,"value":4050},"Provided specific UTM discipline recommendations",{"type":47,"tag":141,"props":4052,"children":4053},{},[4054],{"type":52,"value":4055},"Calculated that 50%+ of upgrades came from in-app pages",{"type":47,"tag":141,"props":4057,"children":4058},{},[4059],{"type":52,"value":4060},"Identified /pricing and /features as high-conversion leverage points",{"type":47,"tag":1057,"props":4062,"children":4064},{"id":4063},"️-models-that-identified-but-stopped-there-40",[4065],{"type":52,"value":4066},"⚠️ Models That Identified But Stopped There (40%)",{"type":47,"tag":48,"props":4068,"children":4069},{},[4070],{"type":52,"value":4071},"These models correctly diagnosed the problem but left users without guidance.",{"type":47,"tag":3840,"props":4073,"children":4075},{"id":4074},"gpt-5",[4076],{"type":52,"value":931},{"type":47,"tag":48,"props":4078,"children":4079},{},[4080,4084,4086,4090],{"type":47,"tag":88,"props":4081,"children":4082},{},[4083],{"type":52,"value":3853},{"type":52,"value":4085}," $0.24 | ",{"type":47,"tag":88,"props":4087,"children":4088},{},[4089],{"type":52,"value":3860},{"type":52,"value":4091}," 163s",{"type":47,"tag":48,"props":4093,"children":4094},{},[4095],{"type":52,"value":4096},"The most thorough diagnostician, but that's where it stopped:",{"type":47,"tag":755,"props":4098,"children":4099},{},[4100],{"type":47,"tag":48,"props":4101,"children":4102},{},[4103],{"type":47,"tag":178,"props":4104,"children":4105},{},[4106],{"type":52,"value":4107},"\"No revenue-expansion events (subscription_upgrade, add_on_purchased) were recorded in the GA4 property over the last 30 days.\"",{"type":47,"tag":48,"props":4109,"children":4110},{},[4111],{"type":52,"value":4112},"GPT-5 provided exhaustive metadata analysis confirming the tracking gap but offered minimal actionable workarounds. If you needed confirmation that something was wrong, this was your model. If you needed help moving forward, you'd be stuck.",{"type":47,"tag":3840,"props":4114,"children":4116},{"id":4115},"gemini-25-flash",[4117],{"type":52,"value":3579},{"type":47,"tag":48,"props":4119,"children":4120},{},[4121,4125,4127,4131],{"type":47,"tag":88,"props":4122,"children":4123},{},[4124],{"type":52,"value":3853},{"type":52,"value":4126}," $0.15 | ",{"type":47,"tag":88,"props":4128,"children":4129},{},[4130],{"type":52,"value":3860},{"type":52,"value":4132}," 27s",{"type":47,"tag":48,"props":4134,"children":4135},{},[4136],{"type":52,"value":4137},"Fast and accurate identification, minimal workarounds:",{"type":47,"tag":755,"props":4139,"children":4140},{},[4141],{"type":47,"tag":48,"props":4142,"children":4143},{},[4144],{"type":47,"tag":178,"props":4145,"children":4146},{},[4147],{"type":52,"value":4148},"\"An attempt to identify top traffic sources... revealed significant data quality issues. All 'sign_up' events were attributed to '(not set)' for both traffic source/medium and landing page.\"",{"type":47,"tag":48,"props":4150,"children":4151},{},[4152],{"type":52,"value":4153},"Good for quick diagnostics but didn't pivot to available data.",{"type":47,"tag":3840,"props":4155,"children":4157},{"id":4156},"deepseek-v32",[4158],{"type":52,"value":3620},{"type":47,"tag":48,"props":4160,"children":4161},{},[4162,4166,4167,4171],{"type":47,"tag":88,"props":4163,"children":4164},{},[4165],{"type":52,"value":3853},{"type":52,"value":4007},{"type":47,"tag":88,"props":4168,"children":4169},{},[4170],{"type":52,"value":3860},{"type":52,"value":4172}," 199s",{"type":47,"tag":48,"props":4174,"children":4175},{},[4176],{"type":52,"value":4177},"Slowest model, accurate diagnosis:",{"type":47,"tag":755,"props":4179,"children":4180},{},[4181],{"type":47,"tag":48,"props":4182,"children":4183},{},[4184],{"type":47,"tag":178,"props":4185,"children":4186},{},[4187],{"type":52,"value":4188},"\"No high-value conversion events (subscription_upgrade, add_on_purchased) were recorded in the last 30 days.\"",{"type":47,"tag":48,"props":4190,"children":4191},{},[4192],{"type":52,"value":4193},"Correctly identified tracking gaps and suggested verification steps, but provided no analysis of what data WAS available.",{"type":47,"tag":3840,"props":4195,"children":4197},{"id":4196},"grok-code-fast-1",[4198],{"type":52,"value":3660},{"type":47,"tag":48,"props":4200,"children":4201},{},[4202,4206,4208,4212],{"type":47,"tag":88,"props":4203,"children":4204},{},[4205],{"type":52,"value":3853},{"type":52,"value":4207}," $0.02 | ",{"type":47,"tag":88,"props":4209,"children":4210},{},[4211],{"type":52,"value":3860},{"type":52,"value":4213}," 28s",{"type":47,"tag":48,"props":4215,"children":4216},{},[4217],{"type":52,"value":4218},"Fast and cheap, correctly flagged the issue but stopped short of extracting value:",{"type":47,"tag":755,"props":4220,"children":4221},{},[4222],{"type":47,"tag":48,"props":4223,"children":4224},{},[4225],{"type":47,"tag":178,"props":4226,"children":4227},{},[4228],{"type":52,"value":4229},"\"Traffic source attribution is entirely missing. All sessions show '(not set)' for source/medium.\"",{"type":47,"tag":48,"props":4231,"children":4232},{},[4233],{"type":52,"value":4234},"Provided a clear diagnosis but no workarounds or analysis of available data.",{"type":47,"tag":1057,"props":4236,"children":4238},{"id":4237},"models-that-hallucinated-or-gave-weak-caveats-30",[4239],{"type":52,"value":4240},"❌ Models That Hallucinated or Gave Weak Caveats (30%)",{"type":47,"tag":48,"props":4242,"children":4243},{},[4244],{"type":52,"value":4245},"These responses ranged from misleading to dangerously wrong.",{"type":47,"tag":3840,"props":4247,"children":4249},{"id":4248},"gemini-25-flash-lite",[4250],{"type":52,"value":3781},{"type":47,"tag":48,"props":4252,"children":4253},{},[4254,4258,4259,4263],{"type":47,"tag":88,"props":4255,"children":4256},{},[4257],{"type":52,"value":3853},{"type":52,"value":4207},{"type":47,"tag":88,"props":4260,"children":4261},{},[4262],{"type":52,"value":3860},{"type":52,"value":4264}," 48s",{"type":47,"tag":48,"props":4266,"children":4267},{},[4268],{"type":52,"value":4269},"The cheapest model. And it showed:",{"type":47,"tag":755,"props":4271,"children":4272},{},[4273],{"type":47,"tag":48,"props":4274,"children":4275},{},[4276],{"type":47,"tag":178,"props":4277,"children":4278},{},[4279],{"type":52,"value":4280},"\"Organic search from Google is the primary driver of high-value users, contributing the most sessions and engagement time.\"",{"type":47,"tag":48,"props":4282,"children":4283},{},[4284,4289],{"type":47,"tag":88,"props":4285,"children":4286},{},[4287],{"type":52,"value":4288},"The problem:",{"type":52,"value":4290}," All traffic data showed \"(not set)\". There was NO organic search data. This model fabricated specific session numbers (15,420 sessions from google/organic) and engagement durations that didn't exist in the dataset.",{"type":47,"tag":48,"props":4292,"children":4293},{},[4294],{"type":52,"value":4295},"If a stakeholder acted on this analysis, they'd be optimizing for phantom traffic sources.",{"type":47,"tag":3840,"props":4297,"children":4299},{"id":4298},"gpt-5-mini",[4300],{"type":52,"value":3741},{"type":47,"tag":48,"props":4302,"children":4303},{},[4304,4308,4310,4314],{"type":47,"tag":88,"props":4305,"children":4306},{},[4307],{"type":52,"value":3853},{"type":52,"value":4309}," $0.05 | ",{"type":47,"tag":88,"props":4311,"children":4312},{},[4313],{"type":52,"value":3860},{"type":52,"value":4315}," 141s",{"type":47,"tag":48,"props":4317,"children":4318},{},[4319],{"type":52,"value":4320},"Technically accurate, completely misleading:",{"type":47,"tag":755,"props":4322,"children":4323},{},[4324],{"type":47,"tag":48,"props":4325,"children":4326},{},[4327],{"type":47,"tag":178,"props":4328,"children":4329},{},[4330],{"type":52,"value":4331},"\"Users whose firstUserSource is recorded as '(direct)' account for ~9,043 of 12,596 subscription_upgrade users (≈71.8%).\"",{"type":47,"tag":48,"props":4333,"children":4334},{},[4335,4339],{"type":47,"tag":88,"props":4336,"children":4337},{},[4338],{"type":52,"value":4288},{"type":52,"value":4340}," While technically true that \"(direct)\" (which includes \"(not set)\") was the largest bucket, presenting this as actionable insight about \"direct traffic\" performance completely misses that this data represents an attribution failure, not a channel insight. Following this advice would lead to meaningless \"direct channel optimization.\"",{"type":47,"tag":3840,"props":4342,"children":4344},{"id":4343},"gemini-3-flash-preview",[4345],{"type":52,"value":3700},{"type":47,"tag":48,"props":4347,"children":4348},{},[4349,4353,4354,4358],{"type":47,"tag":88,"props":4350,"children":4351},{},[4352],{"type":52,"value":3853},{"type":52,"value":4309},{"type":47,"tag":88,"props":4355,"children":4356},{},[4357],{"type":52,"value":3860},{"type":52,"value":4359}," 11s",{"type":47,"tag":48,"props":4361,"children":4362},{},[4363],{"type":52,"value":4364},"The fastest model buried its warning:",{"type":47,"tag":755,"props":4366,"children":4367},{},[4368],{"type":47,"tag":48,"props":4369,"children":4370},{},[4371],{"type":47,"tag":178,"props":4372,"children":4373},{},[4374],{"type":52,"value":4375},"\"A significant portion of traffic is attributed to '(not set)' sources and landing pages, suggesting a need for UTM parameter enforcement.\"",{"type":47,"tag":48,"props":4377,"children":4378},{},[4379],{"type":52,"value":4380},"This caveat appeared well into the analysis, after providing landing page recommendations. The structure implied the analysis was valid with a minor data quality note, rather than flagging that the core question couldn't be answered.",{"type":47,"tag":55,"props":4382,"children":4384},{"id":4383},"what-this-means-for-analytics-products",[4385],{"type":52,"value":4386},"What This Means for Analytics Products",{"type":47,"tag":1057,"props":4388,"children":4390},{"id":4389},"the-speed-cost-quality-tradeoff",[4391],{"type":52,"value":4392},"The Speed-Cost-Quality Tradeoff",{"type":47,"tag":1363,"props":4394,"children":4395},{},[4396,4417],{"type":47,"tag":1367,"props":4397,"children":4398},{},[4399],{"type":47,"tag":1371,"props":4400,"children":4401},{},[4402,4407,4412],{"type":47,"tag":1375,"props":4403,"children":4404},{},[4405],{"type":52,"value":4406},"Metric",{"type":47,"tag":1375,"props":4408,"children":4409},{},[4410],{"type":52,"value":4411},"Winner",{"type":47,"tag":1375,"props":4413,"children":4414},{},[4415],{"type":52,"value":4416},"Loser",{"type":47,"tag":1391,"props":4418,"children":4419},{},[4420,4438,4456,4474,4491],{"type":47,"tag":1371,"props":4421,"children":4422},{},[4423,4428,4433],{"type":47,"tag":1398,"props":4424,"children":4425},{},[4426],{"type":52,"value":4427},"Fastest",{"type":47,"tag":1398,"props":4429,"children":4430},{},[4431],{"type":52,"value":4432},"Gemini 3 Flash (11s)",{"type":47,"tag":1398,"props":4434,"children":4435},{},[4436],{"type":52,"value":4437},"DeepSeek V3.2 (199s)",{"type":47,"tag":1371,"props":4439,"children":4440},{},[4441,4446,4451],{"type":47,"tag":1398,"props":4442,"children":4443},{},[4444],{"type":52,"value":4445},"Cheapest",{"type":47,"tag":1398,"props":4447,"children":4448},{},[4449],{"type":52,"value":4450},"Gemini 2.5 Flash Lite ($0.02)",{"type":47,"tag":1398,"props":4452,"children":4453},{},[4454],{"type":52,"value":4455},"Claude Opus 4.5 ($1.30)",{"type":47,"tag":1371,"props":4457,"children":4458},{},[4459,4464,4469],{"type":47,"tag":1398,"props":4460,"children":4461},{},[4462],{"type":52,"value":4463},"Best Value",{"type":47,"tag":1398,"props":4465,"children":4466},{},[4467],{"type":52,"value":4468},"Grok 4.1 Fast ($0.03, actionable)",{"type":47,"tag":1398,"props":4470,"children":4471},{},[4472],{"type":52,"value":4473},"N/A",{"type":47,"tag":1371,"props":4475,"children":4476},{},[4477,4482,4486],{"type":47,"tag":1398,"props":4478,"children":4479},{},[4480],{"type":52,"value":4481},"Best Analysis",{"type":47,"tag":1398,"props":4483,"children":4484},{},[4485],{"type":52,"value":3416},{"type":47,"tag":1398,"props":4487,"children":4488},{},[4489],{"type":52,"value":4490},"Gemini Flash Lite",{"type":47,"tag":1371,"props":4492,"children":4493},{},[4494,4499,4504],{"type":47,"tag":1398,"props":4495,"children":4496},{},[4497],{"type":52,"value":4498},"Most Thorough",{"type":47,"tag":1398,"props":4500,"children":4501},{},[4502],{"type":52,"value":4503},"GPT-5 (diagnostics)",{"type":47,"tag":1398,"props":4505,"children":4506},{},[4507],{"type":52,"value":4473},{"type":47,"tag":48,"props":4509,"children":4510},{},[4511],{"type":52,"value":4512},"For analytics products, optimizing purely for speed or cost risks deploying models that either:",{"type":47,"tag":522,"props":4514,"children":4515},{},[4516,4526],{"type":47,"tag":141,"props":4517,"children":4518},{},[4519,4524],{"type":47,"tag":88,"props":4520,"children":4521},{},[4522],{"type":52,"value":4523},"Hallucinate insights",{"type":52,"value":4525}," from broken data (dangerous)",{"type":47,"tag":141,"props":4527,"children":4528},{},[4529,4534],{"type":47,"tag":88,"props":4530,"children":4531},{},[4532],{"type":52,"value":4533},"Leave users stuck",{"type":52,"value":4535}," without guidance (frustrating)",{"type":47,"tag":1057,"props":4537,"children":4539},{"id":4538},"the-real-benchmark",[4540],{"type":52,"value":4541},"The Real Benchmark",{"type":47,"tag":48,"props":4543,"children":4544},{},[4545],{"type":52,"value":4546},"Technical accuracy is table stakes; all 10 models passed. What separates useful from dangerous:",{"type":47,"tag":522,"props":4548,"children":4549},{},[4550,4560,4570,4580],{"type":47,"tag":141,"props":4551,"children":4552},{},[4553,4558],{"type":47,"tag":88,"props":4554,"children":4555},{},[4556],{"type":52,"value":4557},"Data quality detection",{"type":52,"value":4559}," - Does the model recognize when data is broken?",{"type":47,"tag":141,"props":4561,"children":4562},{},[4563,4568],{"type":47,"tag":88,"props":4564,"children":4565},{},[4566],{"type":52,"value":4567},"Clear communication",{"type":52,"value":4569}," - Is the issue prominently flagged, not buried?",{"type":47,"tag":141,"props":4571,"children":4572},{},[4573,4578],{"type":47,"tag":88,"props":4574,"children":4575},{},[4576],{"type":52,"value":4577},"Analytical pivot",{"type":52,"value":4579}," - Can it extract value from available data?",{"type":47,"tag":141,"props":4581,"children":4582},{},[4583,4588],{"type":47,"tag":88,"props":4584,"children":4585},{},[4586],{"type":52,"value":4587},"Actionable guidance",{"type":52,"value":4589}," - Does it help users move forward?",{"type":47,"tag":48,"props":4591,"children":4592},{},[4593],{"type":52,"value":4594},"A $0.02 wrong answer costs more than a $1.30 right one.",{"type":47,"tag":1245,"props":4596,"children":4600},{"heading":4597,"icon":1248,"primary-link":1249,"primary-text":4598,"secondary-link":3339,"secondary-text":3340,"subheading":4599},"Ready to try AI-powered analytics?","Start Free Trial","Anamap catches data quality issues before giving recommendations. No hallucinations, no misleading insights.",[],{"type":47,"tag":55,"props":4602,"children":4604},{"id":4603},"methodology",[4605],{"type":52,"value":4606},"Methodology",{"type":47,"tag":48,"props":4608,"children":4609},{},[4610],{"type":47,"tag":88,"props":4611,"children":4612},{},[4613],{"type":52,"value":4614},"Test Setup:",{"type":47,"tag":137,"props":4616,"children":4617},{},[4618,4623,4628,4633],{"type":47,"tag":141,"props":4619,"children":4620},{},[4621],{"type":52,"value":4622},"GA4 property with intentionally broken attribution tracking",{"type":47,"tag":141,"props":4624,"children":4625},{},[4626],{"type":52,"value":4627},"100% of sessions showing \"(not set)\" for source/medium",{"type":47,"tag":141,"props":4629,"children":4630},{},[4631],{"type":52,"value":4632},"Valid conversion events that couldn't be attributed to channels",{"type":47,"tag":141,"props":4634,"children":4635},{},[4636],{"type":52,"value":4637},"Standard marketing ROI question",{"type":47,"tag":48,"props":4639,"children":4640},{},[4641],{"type":47,"tag":88,"props":4642,"children":4643},{},[4644],{"type":52,"value":4645},"Models Tested:",{"type":47,"tag":137,"props":4647,"children":4648},{},[4649,4654,4659,4664,4669],{"type":47,"tag":141,"props":4650,"children":4651},{},[4652],{"type":52,"value":4653},"Anthropic: Claude Opus 4.5, Claude Sonnet 4.5",{"type":47,"tag":141,"props":4655,"children":4656},{},[4657],{"type":52,"value":4658},"OpenAI: GPT-5, GPT-5 Mini",{"type":47,"tag":141,"props":4660,"children":4661},{},[4662],{"type":52,"value":4663},"Google: Gemini 3 Flash Preview, Gemini 2.5 Flash, Gemini 2.5 Flash Lite",{"type":47,"tag":141,"props":4665,"children":4666},{},[4667],{"type":52,"value":4668},"xAI: Grok 4.1 Fast, Grok Code Fast 1",{"type":47,"tag":141,"props":4670,"children":4671},{},[4672],{"type":52,"value":4673},"DeepSeek: V3.2",{"type":47,"tag":48,"props":4675,"children":4676},{},[4677],{"type":47,"tag":88,"props":4678,"children":4679},{},[4680],{"type":52,"value":4681},"Evaluation Criteria:",{"type":47,"tag":137,"props":4683,"children":4684},{},[4685,4690,4695,4700,4705],{"type":47,"tag":141,"props":4686,"children":4687},{},[4688],{"type":52,"value":4689},"Query structure quality (all passed)",{"type":47,"tag":141,"props":4691,"children":4692},{},[4693],{"type":52,"value":4694},"GA4 field name accuracy (97% average, one model at 75%)",{"type":47,"tag":141,"props":4696,"children":4697},{},[4698],{"type":52,"value":4699},"Data quality issue detection",{"type":47,"tag":141,"props":4701,"children":4702},{},[4703],{"type":52,"value":4704},"Actionable workaround provision",{"type":47,"tag":141,"props":4706,"children":4707},{},[4708],{"type":52,"value":4709},"User value delivered",{"type":47,"tag":1047,"props":4711,"children":4712},{},[],{"type":47,"tag":48,"props":4714,"children":4715},{},[4716],{"type":47,"tag":178,"props":4717,"children":4718},{},[4719,4721,4728],{"type":52,"value":4720},"This benchmark was conducted using the ",{"type":47,"tag":67,"props":4722,"children":4725},{"href":4723,"rel":4724},"https://anamaps.com",[71],[4726],{"type":52,"value":4727},"Anamap AI analytics library",{"type":52,"value":4729},", which provides unified analytics querying across GA4 and Amplitude. The test specifically evaluated model behavior when encountering data quality issues, a common real-world scenario that purely technical benchmarks miss.",{"type":47,"tag":55,"props":4731,"children":4732},{"id":1821},[4733],{"type":52,"value":1824},{"type":47,"tag":1057,"props":4735,"children":4737},{"id":4736},"which-llm-is-best-for-analytics",[4738],{"type":52,"value":4739},"Which LLM is best for analytics?",{"type":47,"tag":48,"props":4741,"children":4742},{},[4743,4745,4749,4751,4755],{"type":52,"value":4744},"Based on our benchmark, ",{"type":47,"tag":88,"props":4746,"children":4747},{},[4748],{"type":52,"value":3416},{"type":52,"value":4750}," delivered the best overall analysis, correctly identifying data quality issues while still extracting actionable insights from available data. For budget-conscious users, ",{"type":47,"tag":88,"props":4752,"children":4753},{},[4754],{"type":52,"value":1185},{"type":52,"value":4756}," at $0.03 per query provided the best value with solid analytical capabilities.",{"type":47,"tag":1057,"props":4758,"children":4760},{"id":4759},"do-ai-models-hallucinate-with-analytics-data",[4761],{"type":52,"value":4762},"Do AI models hallucinate with analytics data?",{"type":47,"tag":48,"props":4764,"children":4765},{},[4766],{"type":52,"value":4767},"Yes. In our test with intentionally broken GA4 data (100% attribution failure), 30% of models either hallucinated insights or provided misleadingly framed results. Gemini 2.5 Flash Lite fabricated traffic source data, while GPT-5 Mini presented \"(not set)\" data as actionable \"direct traffic\" insights.",{"type":47,"tag":1057,"props":4769,"children":4771},{"id":4770},"how-much-does-it-cost-to-run-ai-analytics-queries",[4772],{"type":52,"value":4773},"How much does it cost to run AI analytics queries?",{"type":47,"tag":48,"props":4775,"children":4776},{},[4777],{"type":52,"value":4778},"Costs varied significantly across our benchmark:",{"type":47,"tag":137,"props":4780,"children":4781},{},[4782,4792,4802],{"type":47,"tag":141,"props":4783,"children":4784},{},[4785,4790],{"type":47,"tag":88,"props":4786,"children":4787},{},[4788],{"type":52,"value":4789},"Cheapest:",{"type":52,"value":4791}," Gemini 2.5 Flash Lite ($0.02) and Grok Code Fast 1 ($0.02)",{"type":47,"tag":141,"props":4793,"children":4794},{},[4795,4800],{"type":47,"tag":88,"props":4796,"children":4797},{},[4798],{"type":52,"value":4799},"Most Expensive:",{"type":52,"value":4801}," Claude Opus 4.5 ($1.30)",{"type":47,"tag":141,"props":4803,"children":4804},{},[4805,4810],{"type":47,"tag":88,"props":4806,"children":4807},{},[4808],{"type":52,"value":4809},"Best Value:",{"type":52,"value":4811}," Grok 4.1 Fast ($0.03) delivered actionable insights at low cost",{"type":47,"tag":1057,"props":4813,"children":4815},{"id":4814},"can-ai-detect-broken-analytics-tracking",[4816],{"type":52,"value":4817},"Can AI detect broken analytics tracking?",{"type":47,"tag":48,"props":4819,"children":4820},{},[4821],{"type":52,"value":4822},"Our benchmark specifically tested this. Results varied:",{"type":47,"tag":137,"props":4824,"children":4825},{},[4826,4836,4846],{"type":47,"tag":141,"props":4827,"children":4828},{},[4829,4834],{"type":47,"tag":88,"props":4830,"children":4831},{},[4832],{"type":52,"value":4833},"70% detected the issue",{"type":52,"value":4835}," (attribution data showing \"(not set)\")",{"type":47,"tag":141,"props":4837,"children":4838},{},[4839,4844],{"type":47,"tag":88,"props":4840,"children":4841},{},[4842],{"type":52,"value":4843},"30% either missed it",{"type":52,"value":4845}," or buried warnings in their analysis",{"type":47,"tag":141,"props":4847,"children":4848},{},[4849,4853],{"type":47,"tag":88,"props":4850,"children":4851},{},[4852],{"type":52,"value":3253},{"type":52,"value":4854}," both detected AND provided actionable workarounds",{"type":47,"tag":1057,"props":4856,"children":4858},{"id":4857},"which-is-faster-claude-gpt-5-or-gemini",[4859],{"type":52,"value":4860},"Which is faster: Claude, GPT-5, or Gemini?",{"type":47,"tag":48,"props":4862,"children":4863},{},[4864],{"type":52,"value":4865},"Gemini models were fastest in our benchmark:",{"type":47,"tag":137,"props":4867,"children":4868},{},[4869,4879,4889],{"type":47,"tag":141,"props":4870,"children":4871},{},[4872,4877],{"type":47,"tag":88,"props":4873,"children":4874},{},[4875],{"type":52,"value":4876},"Fastest:",{"type":52,"value":4878}," Gemini 3 Flash Preview (11 seconds)",{"type":47,"tag":141,"props":4880,"children":4881},{},[4882,4887],{"type":47,"tag":88,"props":4883,"children":4884},{},[4885],{"type":52,"value":4886},"Slowest:",{"type":52,"value":4888}," DeepSeek V3.2 (199 seconds)",{"type":47,"tag":141,"props":4890,"children":4891},{},[4892],{"type":52,"value":4893},"Claude Opus 4.5 took 96 seconds; GPT-5 took 163 seconds",{"type":47,"tag":1057,"props":4895,"children":4897},{"id":4896},"should-i-use-cheap-ai-models-for-analytics",[4898],{"type":52,"value":4899},"Should I use cheap AI models for analytics?",{"type":47,"tag":48,"props":4901,"children":4902},{},[4903],{"type":52,"value":4904},"It depends on your risk tolerance. The cheapest model (Gemini 2.5 Flash Lite at $0.02) hallucinated data in our test. A wrong answer that leads to misallocated marketing budget costs far more than the price difference between models. For critical decisions, invest in models with better analytical judgment.",{"type":47,"tag":1904,"props":4906,"children":4908},{":faqs":4907},"[{\"question\":\"Which LLM is best for analytics?\",\"answer\":\"Based on our benchmark, Claude Opus 4.5 delivered the best overall analysis, correctly identifying data quality issues while still extracting actionable insights from available data. For budget-conscious users, Grok 4.1 Fast at $0.03 per query provided the best value with solid analytical capabilities.\"},{\"question\":\"Do AI models hallucinate with analytics data?\",\"answer\":\"Yes. In our test with intentionally broken GA4 data (100% attribution failure), 30% of models either hallucinated insights or provided misleadingly framed results. Gemini 2.5 Flash Lite fabricated traffic source data, while GPT-5 Mini presented (not set) data as actionable direct traffic insights.\"},{\"question\":\"How much does it cost to run AI analytics queries?\",\"answer\":\"Costs varied significantly: Cheapest was Gemini 2.5 Flash Lite ($0.02), most expensive was Claude Opus 4.5 ($1.30), and best value was Grok 4.1 Fast ($0.03) which delivered actionable insights at low cost.\"},{\"question\":\"Can AI detect broken analytics tracking?\",\"answer\":\"Our benchmark specifically tested this. 70% detected the issue (attribution data showing not set), 30% either missed it or buried warnings, and only 30% both detected AND provided actionable workarounds.\"},{\"question\":\"Which is faster: Claude, GPT-5, or Gemini?\",\"answer\":\"Gemini models were fastest. Gemini 3 Flash Preview took 11 seconds (fastest), DeepSeek V3.2 took 199 seconds (slowest). Claude Opus 4.5 took 96 seconds; GPT-5 took 163 seconds.\"},{\"question\":\"Should I use cheap AI models for analytics?\",\"answer\":\"It depends on your risk tolerance. The cheapest model (Gemini 2.5 Flash Lite at $0.02) hallucinated data in our test. A wrong answer that leads to misallocated marketing budget costs far more than the price difference between models.\"}]",[],{"type":47,"tag":1047,"props":4910,"children":4911},{},[],{"type":47,"tag":55,"props":4913,"children":4914},{"id":1910},[4915],{"type":52,"value":1913},{"type":47,"tag":137,"props":4917,"children":4918},{},[4919,4929,4939,4948,4957,4967],{"type":47,"tag":141,"props":4920,"children":4921},{},[4922,4927],{"type":47,"tag":67,"props":4923,"children":4924},{"href":912},[4925],{"type":52,"value":4926},"The Best LLM for Analytics in 2026",{"type":52,"value":4928},": Our top recommendations by use case, based on 16 models tested",{"type":47,"tag":141,"props":4930,"children":4931},{},[4932,4937],{"type":47,"tag":67,"props":4933,"children":4934},{"href":1329},[4935],{"type":52,"value":4936},"Round 2: 6 New LLMs, 3 Runs Each",{"type":52,"value":4938},": We tested 6 more models 3 times each. MiniMax M2.5 at $0.06 beat Claude Opus 4.6 at $1.35.",{"type":47,"tag":141,"props":4940,"children":4941},{},[4942,4946],{"type":47,"tag":67,"props":4943,"children":4944},{"href":1040},[4945],{"type":52,"value":1924},{"type":52,"value":4947},": Combined results from all rounds -- 16 models compared in one place",{"type":47,"tag":141,"props":4949,"children":4950},{},[4951,4955],{"type":47,"tag":67,"props":4952,"children":4953},{"href":1962},[4954],{"type":52,"value":1965},{"type":52,"value":4956},": Learn about the data quality issues that trip up both humans and AI",{"type":47,"tag":141,"props":4958,"children":4959},{},[4960,4965],{"type":47,"tag":67,"props":4961,"children":4962},{"href":2200},[4963],{"type":52,"value":4964},"Data Planning for Better Data Quality",{"type":52,"value":4966},": How to set up your analytics implementation to avoid \"(not set)\" nightmares",{"type":47,"tag":141,"props":4968,"children":4969},{},[4970,4974],{"type":47,"tag":67,"props":4971,"children":4972},{"href":603},[4973],{"type":52,"value":604},{"type":52,"value":4975},": Why visual documentation helps teams catch tracking issues faster",{"title":8,"searchDepth":587,"depth":588,"links":4977},[4978,4979,4980,4981,5000,5004,5005,5013],{"id":3225,"depth":587,"text":3228},{"id":3327,"depth":587,"text":3330},{"id":3345,"depth":587,"text":3348},{"id":3824,"depth":587,"text":3827,"children":4982},[4983,4989,4995],{"id":3830,"depth":588,"text":3833,"children":4984},[4985,4987,4988],{"id":3842,"depth":4986,"text":3845},4,{"id":3927,"depth":4986,"text":3930},{"id":3995,"depth":4986,"text":3998},{"id":4063,"depth":588,"text":4066,"children":4990},[4991,4992,4993,4994],{"id":4074,"depth":4986,"text":931},{"id":4115,"depth":4986,"text":3579},{"id":4156,"depth":4986,"text":3620},{"id":4196,"depth":4986,"text":3660},{"id":4237,"depth":588,"text":4240,"children":4996},[4997,4998,4999],{"id":4248,"depth":4986,"text":3781},{"id":4298,"depth":4986,"text":3741},{"id":4343,"depth":4986,"text":3700},{"id":4383,"depth":587,"text":4386,"children":5001},[5002,5003],{"id":4389,"depth":588,"text":4392},{"id":4538,"depth":588,"text":4541},{"id":4603,"depth":587,"text":4606},{"id":1821,"depth":587,"text":1824,"children":5006},[5007,5008,5009,5010,5011,5012],{"id":4736,"depth":588,"text":4739},{"id":4759,"depth":588,"text":4762},{"id":4770,"depth":588,"text":4773},{"id":4814,"depth":588,"text":4817},{"id":4857,"depth":588,"text":4860},{"id":4896,"depth":588,"text":4899},{"id":1910,"depth":587,"text":1913},"content:blog:llm-analytics-benchmark-broken-data.md","blog/llm-analytics-benchmark-broken-data.md","blog/llm-analytics-benchmark-broken-data",{"_path":1040,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":1924,"description":5018,"draft":7,"publicationDate":915,"updatedAt":5019,"image":917,"author":5020,"ogTitle":5021,"ogDescription":5022,"twitterTitle":5023,"twitterDescription":5024,"keywords":5025,"tags":5026,"head":5028,"category":42,"body":5035,"_type":596,"_id":6073,"_source":598,"_file":6074,"_stem":6075,"_extension":601},"26 AI models tested across 3 rounds on real Google Analytics data. The most comprehensive real-world LLM analytics benchmark, updated with each new round. Find the right model for your analytics needs.","2026-06-21",{"id":14,"name":15,"role":16,"twitter":17},"LLM Analytics Benchmark: 26 Models Compared on Real GA4 Data","The definitive leaderboard for AI analytics models. 26 LLMs tested across 58 runs on real Google Analytics data. Updated with every new benchmark round.","🏆 LLM Analytics Benchmark: The Definitive Leaderboard","26 AI models. 58 test runs. Real GA4 data. The most comprehensive LLM analytics comparison, updated with each new round.","LLM analytics leaderboard, best AI for Google Analytics, LLM benchmark comparison, Claude vs GPT-5 vs Gemini vs MiniMax, AI analytics models ranked, LLM cost comparison analytics, AI marketing attribution benchmark, GA4 AI assistant comparison",[3209,25,5027,3213,930,931,932,933,928],"leaderboard",{"meta":5029},[5030,5032,5033,5034],{"name":33,"content":5031},"LLM analytics leaderboard, best AI for Google Analytics, LLM benchmark comparison, Claude vs GPT-5 vs Gemini, AI analytics ranked, LLM cost comparison, AI marketing attribution, GA4 AI comparison",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},{"type":44,"children":5036,"toc":6045},[5037,5043,5053,5058,5107,5113,5116,5122,5128,5136,5141,5149,5200,5208,5214,5222,5227,5234,5286,5294,5300,5308,5313,5320,5369,5377,5383,5389,5479,5485,5636,5642,5682,5686,5692,5735,5741,5812,5818,5823,5846,5851,5857,5860,5881,5885,5889,5905,5911,5916,5922,5927,5933,5938,5944,5983,5989,5994],{"type":47,"tag":55,"props":5038,"children":5040},{"id":5039},"the-most-comprehensive-real-world-llm-analytics-benchmark",[5041],{"type":52,"value":5042},"The Most Comprehensive Real-World LLM Analytics Benchmark",{"type":47,"tag":48,"props":5044,"children":5045},{},[5046,5048],{"type":52,"value":5047},"Most LLM benchmarks test coding puzzles or trivia questions. We test something different: ",{"type":47,"tag":88,"props":5049,"children":5050},{},[5051],{"type":52,"value":5052},"Can this AI model actually help you understand your analytics data?",{"type":47,"tag":48,"props":5054,"children":5055},{},[5056],{"type":52,"value":5057},"This leaderboard combines results from every round of our ongoing benchmark series. Each round tests a new batch of models against connected Google Analytics 4 data. We evaluate not just technical accuracy (API syntax, field names) but analytical judgment: Can the model detect data quality issues, use evidence, provide actionable insights, and help users make better decisions?",{"type":47,"tag":965,"props":5059,"children":5063},{"icon":5060,"title":5061,"type":5062},"mdi-flask","How This Benchmark Works","info",[5064],{"type":47,"tag":137,"props":5065,"children":5066},{},[5067,5077,5087,5097],{"type":47,"tag":141,"props":5068,"children":5069},{},[5070,5075],{"type":47,"tag":88,"props":5071,"children":5072},{},[5073],{"type":52,"value":5074},"Real GA4 API access",{"type":52,"value":5076}," -- every round uses connected analytics data and the same query runner.",{"type":47,"tag":141,"props":5078,"children":5079},{},[5080,5085],{"type":47,"tag":88,"props":5081,"children":5082},{},[5083],{"type":52,"value":5084},"Different analytics tasks",{"type":52,"value":5086}," -- broken attribution, consistency testing, and synthetic-data detection.",{"type":47,"tag":141,"props":5088,"children":5089},{},[5090,5095],{"type":47,"tag":88,"props":5091,"children":5092},{},[5093],{"type":52,"value":5094},"Holistic evaluation",{"type":52,"value":5096}," -- technical accuracy, data quality detection, evidence quality, analytical depth, and actionable guidance.",{"type":47,"tag":141,"props":5098,"children":5099},{},[5100,5105],{"type":47,"tag":88,"props":5101,"children":5102},{},[5103],{"type":52,"value":5104},"Multi-run testing",{"type":52,"value":5106}," -- newer rounds test each model 3 times to measure consistency",{"type":47,"tag":55,"props":5108,"children":5110},{"id":5109},"combined-leaderboard",[5111],{"type":52,"value":5112},"Combined Leaderboard",{"type":47,"tag":1522,"props":5114,"children":5115},{},[],{"type":47,"tag":55,"props":5117,"children":5119},{"id":5118},"round-by-round-results",[5120],{"type":52,"value":5121},"Round-by-Round Results",{"type":47,"tag":1057,"props":5123,"children":5125},{"id":5124},"round-3-synthetic-data-detection-june-2026",[5126],{"type":52,"value":5127},"Round 3: Synthetic Data Detection (June 2026)",{"type":47,"tag":48,"props":5129,"children":5130},{},[5131],{"type":47,"tag":88,"props":5132,"children":5133},{},[5134],{"type":52,"value":5135},"10 models, 3 runs each, 30 total test runs",{"type":47,"tag":48,"props":5137,"children":5138},{},[5139],{"type":52,"value":5140},"The third round tested whether newer models could determine if a connected GA4 dataset was real production data, synthetic demo data, or inconclusive. We used the same analytics runner but changed the prompt: instead of asking for marketing recommendations, we asked models to show evidence for data provenance.",{"type":47,"tag":48,"props":5142,"children":5143},{},[5144],{"type":47,"tag":88,"props":5145,"children":5146},{},[5147],{"type":52,"value":5148},"Key findings:",{"type":47,"tag":137,"props":5150,"children":5151},{},[5152,5161,5171,5180,5190],{"type":47,"tag":141,"props":5153,"children":5154},{},[5155,5159],{"type":47,"tag":88,"props":5156,"children":5157},{},[5158],{"type":52,"value":1107},{"type":52,"value":5160}," produced the best evidence-backed synthetic-data audit",{"type":47,"tag":141,"props":5162,"children":5163},{},[5164,5169],{"type":47,"tag":88,"props":5165,"children":5166},{},[5167],{"type":52,"value":5168},"Grok 4.3",{"type":52,"value":5170}," was the fastest successful classifier, but made no data requests",{"type":47,"tag":141,"props":5172,"children":5173},{},[5174,5178],{"type":47,"tag":88,"props":5175,"children":5176},{},[5177],{"type":52,"value":1205},{"type":52,"value":5179}," gave a nuanced low-cost answer, though it was slow",{"type":47,"tag":141,"props":5181,"children":5182},{},[5183,5188],{"type":47,"tag":88,"props":5184,"children":5185},{},[5186],{"type":52,"value":5187},"GPT-5.5",{"type":52,"value":5189}," correctly identified the dataset as synthetic but had the lowest field-accuracy score among successful Round 3 models",{"type":47,"tag":141,"props":5191,"children":5192},{},[5193,5198],{"type":47,"tag":88,"props":5194,"children":5195},{},[5196],{"type":52,"value":5197},"9 of 10 replacement models",{"type":52,"value":5199}," completed successfully after replacing Claude Fable 5, which was unavailable after a U.S. export-control directive, with GPT-5.5",{"type":47,"tag":48,"props":5201,"children":5202},{},[5203],{"type":47,"tag":67,"props":5204,"children":5205},{"href":1344},[5206],{"type":52,"value":5207},"Read the full Round 3 analysis",{"type":47,"tag":1057,"props":5209,"children":5211},{"id":5210},"round-2-consistency-test-february-2026",[5212],{"type":52,"value":5213},"Round 2: Consistency Test (February 2026)",{"type":47,"tag":48,"props":5215,"children":5216},{},[5217],{"type":47,"tag":88,"props":5218,"children":5219},{},[5220],{"type":52,"value":5221},"6 models, 3 runs each, 18 total test runs",{"type":47,"tag":48,"props":5223,"children":5224},{},[5225],{"type":52,"value":5226},"The second round focused on consistency, testing whether AI models deliver reliable results across multiple runs. We also expanded to include models from Chinese AI labs and a stealth OpenRouter release.",{"type":47,"tag":48,"props":5228,"children":5229},{},[5230],{"type":47,"tag":88,"props":5231,"children":5232},{},[5233],{"type":52,"value":5148},{"type":47,"tag":137,"props":5235,"children":5236},{},[5237,5246,5256,5266,5276],{"type":47,"tag":141,"props":5238,"children":5239},{},[5240,5244],{"type":47,"tag":88,"props":5241,"children":5242},{},[5243],{"type":52,"value":1070},{"type":52,"value":5245}," dominated on every efficiency metric -- fastest, cheapest, excellent quality",{"type":47,"tag":141,"props":5247,"children":5248},{},[5249,5254],{"type":47,"tag":88,"props":5250,"children":5251},{},[5252],{"type":52,"value":5253},"5 of 6 models",{"type":52,"value":5255}," achieved excellent quality, a significant improvement over Round 1's quality variance",{"type":47,"tag":141,"props":5257,"children":5258},{},[5259,5264],{"type":47,"tag":88,"props":5260,"children":5261},{},[5262],{"type":52,"value":5263},"Aurora Alpha",{"type":52,"value":5265}," (stealth OpenRouter release) failed all 3 runs due to context window limitations",{"type":47,"tag":141,"props":5267,"children":5268},{},[5269,5274],{"type":47,"tag":88,"props":5270,"children":5271},{},[5272],{"type":52,"value":5273},"Quality was consistent",{"type":52,"value":5275}," -- 14 of 15 successful runs scored \"excellent\"",{"type":47,"tag":141,"props":5277,"children":5278},{},[5279,5284],{"type":47,"tag":88,"props":5280,"children":5281},{},[5282],{"type":52,"value":5283},"Speed varied significantly",{"type":52,"value":5285}," -- GLM 5 ranged from 145s to 275s across runs",{"type":47,"tag":48,"props":5287,"children":5288},{},[5289],{"type":47,"tag":67,"props":5290,"children":5291},{"href":1329},[5292],{"type":52,"value":5293},"Read the full Round 2 analysis",{"type":47,"tag":1057,"props":5295,"children":5297},{"id":5296},"round-1-the-broken-data-test-january-2026",[5298],{"type":52,"value":5299},"Round 1: The Broken Data Test (January 2026)",{"type":47,"tag":48,"props":5301,"children":5302},{},[5303],{"type":47,"tag":88,"props":5304,"children":5305},{},[5306],{"type":52,"value":5307},"10 models, 1 run each, 10 total test runs",{"type":47,"tag":48,"props":5309,"children":5310},{},[5311],{"type":52,"value":5312},"The original benchmark tested how leading AI models handle a common real-world scenario: broken analytics data. All traffic attribution showed as \"(not set)\" with zero conversion tracking.",{"type":47,"tag":48,"props":5314,"children":5315},{},[5316],{"type":47,"tag":88,"props":5317,"children":5318},{},[5319],{"type":52,"value":5148},{"type":47,"tag":137,"props":5321,"children":5322},{},[5323,5332,5341,5351,5360],{"type":47,"tag":141,"props":5324,"children":5325},{},[5326,5330],{"type":47,"tag":88,"props":5327,"children":5328},{},[5329],{"type":52,"value":3243},{"type":52,"value":5331}," achieved perfect API syntax -- technical accuracy is table stakes",{"type":47,"tag":141,"props":5333,"children":5334},{},[5335,5339],{"type":47,"tag":88,"props":5336,"children":5337},{},[5338],{"type":52,"value":3253},{"type":52,"value":5340}," provided actionable insights despite broken data",{"type":47,"tag":141,"props":5342,"children":5343},{},[5344,5349],{"type":47,"tag":88,"props":5345,"children":5346},{},[5347],{"type":52,"value":5348},"30% hallucinated",{"type":52,"value":5350}," -- fabricating traffic source data or presenting broken data as insights",{"type":47,"tag":141,"props":5352,"children":5353},{},[5354,5358],{"type":47,"tag":88,"props":5355,"children":5356},{},[5357],{"type":52,"value":3416},{"type":52,"value":5359}," delivered the best analysis with workarounds and next steps",{"type":47,"tag":141,"props":5361,"children":5362},{},[5363,5367],{"type":47,"tag":88,"props":5364,"children":5365},{},[5366],{"type":52,"value":1185},{"type":52,"value":5368}," was the Round 1 best value at $0.03 with solid analysis",{"type":47,"tag":48,"props":5370,"children":5371},{},[5372],{"type":47,"tag":67,"props":5373,"children":5374},{"href":1313},[5375],{"type":52,"value":5376},"Read the full Round 1 analysis",{"type":47,"tag":55,"props":5378,"children":5380},{"id":5379},"how-to-use-this-data",[5381],{"type":52,"value":5382},"How to Use This Data",{"type":47,"tag":1057,"props":5384,"children":5386},{"id":5385},"choosing-by-budget",[5387],{"type":52,"value":5388},"Choosing by Budget",{"type":47,"tag":1363,"props":5390,"children":5391},{},[5392,5413],{"type":47,"tag":1367,"props":5393,"children":5394},{},[5395],{"type":47,"tag":1371,"props":5396,"children":5397},{},[5398,5403,5408],{"type":47,"tag":1375,"props":5399,"children":5400},{},[5401],{"type":52,"value":5402},"Budget",{"type":47,"tag":1375,"props":5404,"children":5405},{},[5406],{"type":52,"value":5407},"Best Choice",{"type":47,"tag":1375,"props":5409,"children":5410},{},[5411],{"type":52,"value":5412},"Why",{"type":47,"tag":1391,"props":5414,"children":5415},{},[5416,5437,5458],{"type":47,"tag":1371,"props":5417,"children":5418},{},[5419,5427,5432],{"type":47,"tag":1398,"props":5420,"children":5421},{},[5422],{"type":47,"tag":88,"props":5423,"children":5424},{},[5425],{"type":52,"value":5426},"Under $0.05/query",{"type":47,"tag":1398,"props":5428,"children":5429},{},[5430],{"type":52,"value":5431},"Grok 4.1 Fast ($0.03, R1) or MiniMax M2.5 ($0.02/run, R2)",{"type":47,"tag":1398,"props":5433,"children":5434},{},[5435],{"type":52,"value":5436},"Both delivered excellent quality at rock-bottom prices",{"type":47,"tag":1371,"props":5438,"children":5439},{},[5440,5448,5453],{"type":47,"tag":1398,"props":5441,"children":5442},{},[5443],{"type":47,"tag":88,"props":5444,"children":5445},{},[5446],{"type":52,"value":5447},"Under $0.25/query",{"type":47,"tag":1398,"props":5449,"children":5450},{},[5451],{"type":52,"value":5452},"Kimi K2.5 ($0.07, R2), Qwen3.7 Max ($0.12, R3), or Gemini 3.5 Flash ($0.23, R3)",{"type":47,"tag":1398,"props":5454,"children":5455},{},[5456],{"type":52,"value":5457},"Strong analysis with good speed and evidence",{"type":47,"tag":1371,"props":5459,"children":5460},{},[5461,5469,5474],{"type":47,"tag":1398,"props":5462,"children":5463},{},[5464],{"type":47,"tag":88,"props":5465,"children":5466},{},[5467],{"type":52,"value":5468},"No budget limit",{"type":47,"tag":1398,"props":5470,"children":5471},{},[5472],{"type":52,"value":5473},"Claude Opus 4.8 ($0.81, R3), GPT-5.5 ($1.45, R3), or Claude Opus 4.6 ($1.35, R2)",{"type":47,"tag":1398,"props":5475,"children":5476},{},[5477],{"type":52,"value":5478},"Most comprehensive, thorough investigation",{"type":47,"tag":1057,"props":5480,"children":5482},{"id":5481},"choosing-by-use-case",[5483],{"type":52,"value":5484},"Choosing by Use Case",{"type":47,"tag":1363,"props":5486,"children":5487},{},[5488,5509],{"type":47,"tag":1367,"props":5489,"children":5490},{},[5491],{"type":47,"tag":1371,"props":5492,"children":5493},{},[5494,5499,5504],{"type":47,"tag":1375,"props":5495,"children":5496},{},[5497],{"type":52,"value":5498},"Use Case",{"type":47,"tag":1375,"props":5500,"children":5501},{},[5502],{"type":52,"value":5503},"Recommended Model",{"type":47,"tag":1375,"props":5505,"children":5506},{},[5507],{"type":52,"value":5508},"Reason",{"type":47,"tag":1391,"props":5510,"children":5511},{},[5512,5532,5553,5573,5594,5615],{"type":47,"tag":1371,"props":5513,"children":5514},{},[5515,5523,5527],{"type":47,"tag":1398,"props":5516,"children":5517},{},[5518],{"type":47,"tag":88,"props":5519,"children":5520},{},[5521],{"type":52,"value":5522},"Daily automated queries",{"type":47,"tag":1398,"props":5524,"children":5525},{},[5526],{"type":52,"value":1070},{"type":47,"tag":1398,"props":5528,"children":5529},{},[5530],{"type":52,"value":5531},"Cheapest + fastest at excellent quality",{"type":47,"tag":1371,"props":5533,"children":5534},{},[5535,5543,5548],{"type":47,"tag":1398,"props":5536,"children":5537},{},[5538],{"type":47,"tag":88,"props":5539,"children":5540},{},[5541],{"type":52,"value":5542},"Executive dashboards",{"type":47,"tag":1398,"props":5544,"children":5545},{},[5546],{"type":52,"value":5547},"Claude Opus 4.8 or Claude Opus 4.6",{"type":47,"tag":1398,"props":5549,"children":5550},{},[5551],{"type":52,"value":5552},"Most thorough, catches nuances",{"type":47,"tag":1371,"props":5554,"children":5555},{},[5556,5564,5568],{"type":47,"tag":1398,"props":5557,"children":5558},{},[5559],{"type":47,"tag":88,"props":5560,"children":5561},{},[5562],{"type":52,"value":5563},"Quick diagnostics",{"type":47,"tag":1398,"props":5565,"children":5566},{},[5567],{"type":52,"value":3700},{"type":47,"tag":1398,"props":5569,"children":5570},{},[5571],{"type":52,"value":5572},"11s response time",{"type":47,"tag":1371,"props":5574,"children":5575},{},[5576,5584,5589],{"type":47,"tag":1398,"props":5577,"children":5578},{},[5579],{"type":47,"tag":88,"props":5580,"children":5581},{},[5582],{"type":52,"value":5583},"Budget analytics teams",{"type":47,"tag":1398,"props":5585,"children":5586},{},[5587],{"type":52,"value":5588},"Grok 4.1 Fast or Kimi K2.5",{"type":47,"tag":1398,"props":5590,"children":5591},{},[5592],{"type":52,"value":5593},"Excellent analysis under $0.10",{"type":47,"tag":1371,"props":5595,"children":5596},{},[5597,5605,5610],{"type":47,"tag":1398,"props":5598,"children":5599},{},[5600],{"type":47,"tag":88,"props":5601,"children":5602},{},[5603],{"type":52,"value":5604},"Synthetic-data audits",{"type":47,"tag":1398,"props":5606,"children":5607},{},[5608],{"type":52,"value":5609},"Gemini 3.5 Flash or Qwen3.7 Max",{"type":47,"tag":1398,"props":5611,"children":5612},{},[5613],{"type":52,"value":5614},"Best evidence-backed Round 3 results",{"type":47,"tag":1371,"props":5616,"children":5617},{},[5618,5626,5631],{"type":47,"tag":1398,"props":5619,"children":5620},{},[5621],{"type":47,"tag":88,"props":5622,"children":5623},{},[5624],{"type":52,"value":5625},"Data quality audits",{"type":47,"tag":1398,"props":5627,"children":5628},{},[5629],{"type":52,"value":5630},"Claude Opus 4.8, Claude Opus 4.6, or Gemini 3.5 Flash",{"type":47,"tag":1398,"props":5632,"children":5633},{},[5634],{"type":52,"value":5635},"Best at finding and explaining issues",{"type":47,"tag":1057,"props":5637,"children":5639},{"id":5638},"models-to-avoid",[5640],{"type":52,"value":5641},"Models to Avoid",{"type":47,"tag":137,"props":5643,"children":5644},{},[5645,5654,5663,5672],{"type":47,"tag":141,"props":5646,"children":5647},{},[5648,5652],{"type":47,"tag":88,"props":5649,"children":5650},{},[5651],{"type":52,"value":3781},{"type":52,"value":5653}," -- hallucinated traffic source data in our test. A wrong answer is worse than no answer.",{"type":47,"tag":141,"props":5655,"children":5656},{},[5657,5661],{"type":47,"tag":88,"props":5658,"children":5659},{},[5660],{"type":52,"value":3741},{"type":52,"value":5662}," -- presented broken data as actionable \"direct traffic\" insights without adequate caveats.",{"type":47,"tag":141,"props":5664,"children":5665},{},[5666,5670],{"type":47,"tag":88,"props":5667,"children":5668},{},[5669],{"type":52,"value":5263},{"type":52,"value":5671}," -- failed to complete analysis due to context window limitations.",{"type":47,"tag":141,"props":5673,"children":5674},{},[5675,5680],{"type":47,"tag":88,"props":5676,"children":5677},{},[5678],{"type":52,"value":5679},"Claude Fable 5",{"type":52,"value":5681}," -- failed all 3 Round 3 attempts through OpenRouter after access to Fable 5 was disabled following a U.S. export-control directive, so it was replaced with GPT-5.5 for the final analysis.",{"type":47,"tag":55,"props":5683,"children":5684},{"id":4603},[5685],{"type":52,"value":4606},{"type":47,"tag":1057,"props":5687,"children":5689},{"id":5688},"test-environment",[5690],{"type":52,"value":5691},"Test Environment",{"type":47,"tag":137,"props":5693,"children":5694},{},[5695,5705,5715,5725],{"type":47,"tag":141,"props":5696,"children":5697},{},[5698,5703],{"type":47,"tag":88,"props":5699,"children":5700},{},[5701],{"type":52,"value":5702},"Data source:",{"type":52,"value":5704}," Real Google Analytics 4 property",{"type":47,"tag":141,"props":5706,"children":5707},{},[5708,5713],{"type":47,"tag":88,"props":5709,"children":5710},{},[5711],{"type":52,"value":5712},"Data condition:",{"type":52,"value":5714}," Intentionally broken attribution tracking in Rounds 1-2; synthetic demo-data provenance audit in Round 3",{"type":47,"tag":141,"props":5716,"children":5717},{},[5718,5723],{"type":47,"tag":88,"props":5719,"children":5720},{},[5721],{"type":52,"value":5722},"Queries:",{"type":52,"value":5724}," Rounds 1-2 used a marketing attribution query. Round 3 used a synthetic-data provenance prompt.",{"type":47,"tag":141,"props":5726,"children":5727},{},[5728,5733],{"type":47,"tag":88,"props":5729,"children":5730},{},[5731],{"type":52,"value":5732},"Valid conversion events:",{"type":52,"value":5734}," sign_up, subscription_upgrade, add_on_purchased (properly tracked)",{"type":47,"tag":1057,"props":5736,"children":5738},{"id":5737},"what-we-evaluate",[5739],{"type":52,"value":5740},"What We Evaluate",{"type":47,"tag":522,"props":5742,"children":5743},{},[5744,5754,5764,5773,5783,5793,5802],{"type":47,"tag":141,"props":5745,"children":5746},{},[5747,5752],{"type":47,"tag":88,"props":5748,"children":5749},{},[5750],{"type":52,"value":5751},"Technical accuracy",{"type":52,"value":5753}," -- Valid GA4 field names, correct API syntax, proper query structure",{"type":47,"tag":141,"props":5755,"children":5756},{},[5757,5762],{"type":47,"tag":88,"props":5758,"children":5759},{},[5760],{"type":52,"value":5761},"Accuracy score",{"type":52,"value":5763}," (0-100) -- How accurately the model uses real GA4 dimensions and metrics",{"type":47,"tag":141,"props":5765,"children":5766},{},[5767,5771],{"type":47,"tag":88,"props":5768,"children":5769},{},[5770],{"type":52,"value":4557},{"type":52,"value":5772}," -- Does the model identify attribution tracking issues?",{"type":47,"tag":141,"props":5774,"children":5775},{},[5776,5781],{"type":47,"tag":88,"props":5777,"children":5778},{},[5779],{"type":52,"value":5780},"Analytical depth",{"type":52,"value":5782}," -- How many data requests? How thorough is the investigation?",{"type":47,"tag":141,"props":5784,"children":5785},{},[5786,5791],{"type":47,"tag":88,"props":5787,"children":5788},{},[5789],{"type":52,"value":5790},"Actionable output",{"type":52,"value":5792}," -- Does the model provide guidance, workarounds, and next steps?",{"type":47,"tag":141,"props":5794,"children":5795},{},[5796,5800],{"type":47,"tag":88,"props":5797,"children":5798},{},[5799],{"type":52,"value":1501},{"type":52,"value":5801}," (multi-run rounds) -- Does the model deliver similar quality every time?",{"type":47,"tag":141,"props":5803,"children":5804},{},[5805,5810],{"type":47,"tag":88,"props":5806,"children":5807},{},[5808],{"type":52,"value":5809},"Evidence quality",{"type":52,"value":5811}," (Round 3) -- Does the model actually inspect connected data before judging whether a dataset is real or synthetic?",{"type":47,"tag":1057,"props":5813,"children":5815},{"id":5814},"how-models-are-ranked",[5816],{"type":52,"value":5817},"How Models Are Ranked",{"type":47,"tag":48,"props":5819,"children":5820},{},[5821],{"type":52,"value":5822},"The leaderboard ranks models by a composite score weighing:",{"type":47,"tag":137,"props":5824,"children":5825},{},[5826,5831,5836,5841],{"type":47,"tag":141,"props":5827,"children":5828},{},[5829],{"type":52,"value":5830},"Quality rating (40%) -- overall analytical value delivered",{"type":47,"tag":141,"props":5832,"children":5833},{},[5834],{"type":52,"value":5835},"Cost efficiency (25%) -- $/1K tokens",{"type":47,"tag":141,"props":5837,"children":5838},{},[5839],{"type":52,"value":5840},"Accuracy score (20%) -- GA4 field name correctness",{"type":47,"tag":141,"props":5842,"children":5843},{},[5844],{"type":52,"value":5845},"Speed (15%) -- average response time",{"type":47,"tag":48,"props":5847,"children":5848},{},[5849],{"type":52,"value":5850},"Models that hallucinate or provide misleading results are ranked below models that correctly identify limitations, regardless of other metrics.",{"type":47,"tag":1245,"props":5852,"children":5856},{"heading":5853,"icon":5854,"primary-link":1249,"primary-text":1250,"secondary-link":3339,"secondary-text":3340,"subheading":5855},"Want AI analytics that works?","mdi-shield-check-outline","Anamap uses rigorously benchmarked models to deliver reliable analytics insights. No hallucinations, no misleading results.",[],{"type":47,"tag":1047,"props":5858,"children":5859},{},[],{"type":47,"tag":48,"props":5861,"children":5862},{},[5863],{"type":47,"tag":178,"props":5864,"children":5865},{},[5866,5868,5873,5875],{"type":52,"value":5867},"This leaderboard is updated with each new benchmark round. All testing is conducted using the ",{"type":47,"tag":67,"props":5869,"children":5871},{"href":4723,"rel":5870},[71],[5872],{"type":52,"value":4727},{"type":52,"value":5874},". Want to suggest a model for the next round? ",{"type":47,"tag":67,"props":5876,"children":5878},{"href":5877},"/contact",[5879],{"type":52,"value":5880},"Let us know.",{"type":47,"tag":55,"props":5882,"children":5883},{"id":1821},[5884],{"type":52,"value":1824},{"type":47,"tag":1057,"props":5886,"children":5887},{"id":1827},[5888],{"type":52,"value":1830},{"type":47,"tag":48,"props":5890,"children":5891},{},[5892,5894,5898,5900,5904],{"type":52,"value":5893},"Based on our benchmark of 26 models across 58 test runs, ",{"type":47,"tag":88,"props":5895,"children":5896},{},[5897],{"type":52,"value":1070},{"type":52,"value":5899}," still offers the best combination of quality, speed, and cost for everyday marketing analytics at about $0.02 per query. For synthetic-data audits, ",{"type":47,"tag":88,"props":5901,"children":5902},{},[5903],{"type":52,"value":1107},{"type":52,"value":1847},{"type":47,"tag":1057,"props":5906,"children":5908},{"id":5907},"how-often-is-this-leaderboard-updated",[5909],{"type":52,"value":5910},"How often is this leaderboard updated?",{"type":47,"tag":48,"props":5912,"children":5913},{},[5914],{"type":52,"value":5915},"We run new benchmark rounds periodically, testing fresh batches of models as they release. Each round is documented in a detailed blog post, and results are added to this combined leaderboard. The current data reflects 3 rounds from January, February, and June 2026.",{"type":47,"tag":1057,"props":5917,"children":5919},{"id":5918},"why-test-on-broken-analytics-data",[5920],{"type":52,"value":5921},"Why test on broken analytics data?",{"type":47,"tag":48,"props":5923,"children":5924},{},[5925],{"type":52,"value":5926},"Broken attribution is one of the most common real-world analytics problems. Testing on clean, well-structured data only measures technical capability. Our benchmark measures analytical judgment: can the AI detect problems, communicate them clearly, and still extract value?",{"type":47,"tag":1057,"props":5928,"children":5930},{"id":5929},"can-i-trust-cheap-ai-models-for-analytics",[5931],{"type":52,"value":5932},"Can I trust cheap AI models for analytics?",{"type":47,"tag":48,"props":5934,"children":5935},{},[5936],{"type":52,"value":5937},"Yes, with caveats. Our Round 2 results show that MiniMax M2.5 and Kimi K2.5 delivered excellent quality at very low cost. Round 3 also showed low-cost wins from Gemini 3.1 Flash Lite and MiniMax M3. However, Round 1 showed that a cheap model can still hallucinate data. Always validate that a model handles your edge cases before relying on it for production analytics.",{"type":47,"tag":1057,"props":5939,"children":5941},{"id":5940},"which-ai-providers-make-the-best-analytics-models",[5942],{"type":52,"value":5943},"Which AI providers make the best analytics models?",{"type":47,"tag":48,"props":5945,"children":5946},{},[5947,5949,5954,5956,5960,5962,5967,5969,5974,5976,5981],{"type":52,"value":5948},"Based on our data: ",{"type":47,"tag":88,"props":5950,"children":5951},{},[5952],{"type":52,"value":5953},"Anthropic",{"type":52,"value":5955}," (Claude) leads on analytical depth but is often expensive. ",{"type":47,"tag":88,"props":5957,"children":5958},{},[5959],{"type":52,"value":933},{"type":52,"value":5961}," remains strong on cost efficiency. ",{"type":47,"tag":88,"props":5963,"children":5964},{},[5965],{"type":52,"value":5966},"Google",{"type":52,"value":5968}," improved in Round 3, with Gemini 3.5 Flash producing the best evidence-backed synthetic-data audit. ",{"type":47,"tag":88,"props":5970,"children":5971},{},[5972],{"type":52,"value":5973},"xAI",{"type":52,"value":5975}," is consistently fast. ",{"type":47,"tag":88,"props":5977,"children":5978},{},[5979],{"type":52,"value":5980},"OpenAI",{"type":52,"value":5982}," results are mixed: GPT-5.5 was useful in Round 3, while GPT-5 Mini was misleading in Round 1.",{"type":47,"tag":1057,"props":5984,"children":5986},{"id":5985},"how-do-you-measure-llm-accuracy-in-analytics",[5987],{"type":52,"value":5988},"How do you measure LLM accuracy in analytics?",{"type":47,"tag":48,"props":5990,"children":5991},{},[5992],{"type":52,"value":5993},"We track an accuracy score (0-100) that measures how correctly each model uses valid GA4 API field names. A score of 100 means every dimension and metric used was valid. We also evaluate whether models fabricate data, present broken data as reliable insights, or judge data provenance without sufficient evidence.",{"type":47,"tag":1904,"props":5995,"children":5997},{":faqs":5996},"[{\"question\":\"What is the best LLM for Google Analytics?\",\"answer\":\"Based on our benchmark of 26 models across 58 test runs, MiniMax M2.5 still offers the best combination of quality, speed, and cost for everyday marketing analytics. For synthetic-data audits, Gemini 3.5 Flash produced the best evidence-backed Round 3 result.\"},{\"question\":\"How often is this leaderboard updated?\",\"answer\":\"We run new benchmark rounds periodically, testing fresh batches of models as they release. Each round is documented in a detailed blog post, and results are added to this combined leaderboard. The current data reflects 3 rounds from January, February, and June 2026.\"},{\"question\":\"Why test on broken analytics data?\",\"answer\":\"Broken attribution is one of the most common real-world analytics problems. Testing on clean data only measures technical capability. Our benchmark measures analytical judgment: can the AI detect problems, communicate them clearly, and still extract value.\"},{\"question\":\"Can I trust cheap AI models for analytics?\",\"answer\":\"Yes, with caveats. MiniMax M2.5 and Kimi K2.5 delivered excellent Round 2 quality at very low cost, and Round 3 added low-cost wins from Gemini 3.1 Flash Lite and MiniMax M3. However, Round 1 showed that a cheap model can still hallucinate data. Always validate edge case handling before relying on a model for production.\"},{\"question\":\"Which AI providers make the best analytics models?\",\"answer\":\"Anthropic leads on analytical depth but is often expensive. MiniMax remains strong on cost efficiency. Google improved in Round 3 with Gemini 3.5 Flash producing the best evidence-backed synthetic-data audit. xAI is consistently fast. OpenAI results are mixed: GPT-5.5 was useful in Round 3, while GPT-5 Mini was misleading in Round 1.\"},{\"question\":\"How do you measure LLM accuracy in analytics?\",\"answer\":\"We track an accuracy score from 0 to 100 that measures how correctly each model uses valid GA4 API field names. We also evaluate whether models fabricate data, present broken data as reliable insights, or judge data provenance without sufficient evidence.\"}]",[5998,6001,6005],{"type":47,"tag":1047,"props":5999,"children":6000},{},[],{"type":47,"tag":55,"props":6002,"children":6003},{"id":1910},[6004],{"type":52,"value":1913},{"type":47,"tag":137,"props":6006,"children":6007},{},[6008,6017,6025,6035],{"type":47,"tag":141,"props":6009,"children":6010},{},[6011,6015],{"type":47,"tag":67,"props":6012,"children":6013},{"href":912},[6014],{"type":52,"value":4926},{"type":52,"value":6016},": Our top recommendations by use case, based on all benchmark data",{"type":47,"tag":141,"props":6018,"children":6019},{},[6020,6024],{"type":47,"tag":67,"props":6021,"children":6022},{"href":1344},[6023],{"type":52,"value":1934},{"type":52,"value":1936},{"type":47,"tag":141,"props":6026,"children":6027},{},[6028,6033],{"type":47,"tag":67,"props":6029,"children":6030},{"href":1313},[6031],{"type":52,"value":6032},"Round 1: I Benchmarked 10 LLMs on Broken Analytics Data",{"type":52,"value":6034},": The original benchmark testing 10 established models",{"type":47,"tag":141,"props":6036,"children":6037},{},[6038,6043],{"type":47,"tag":67,"props":6039,"children":6040},{"href":1329},[6041],{"type":52,"value":6042},"Round 2: The Consistency Test — 6 Models, 3 Runs Each",{"type":52,"value":6044},": Testing newer models for consistency and cost efficiency",{"title":8,"searchDepth":587,"depth":588,"links":6046},[6047,6048,6049,6054,6059,6064,6072],{"id":5039,"depth":587,"text":5042},{"id":5109,"depth":587,"text":5112},{"id":5118,"depth":587,"text":5121,"children":6050},[6051,6052,6053],{"id":5124,"depth":588,"text":5127},{"id":5210,"depth":588,"text":5213},{"id":5296,"depth":588,"text":5299},{"id":5379,"depth":587,"text":5382,"children":6055},[6056,6057,6058],{"id":5385,"depth":588,"text":5388},{"id":5481,"depth":588,"text":5484},{"id":5638,"depth":588,"text":5641},{"id":4603,"depth":587,"text":4606,"children":6060},[6061,6062,6063],{"id":5688,"depth":588,"text":5691},{"id":5737,"depth":588,"text":5740},{"id":5814,"depth":588,"text":5817},{"id":1821,"depth":587,"text":1824,"children":6065},[6066,6067,6068,6069,6070,6071],{"id":1827,"depth":588,"text":1830},{"id":5907,"depth":588,"text":5910},{"id":5918,"depth":588,"text":5921},{"id":5929,"depth":588,"text":5932},{"id":5940,"depth":588,"text":5943},{"id":5985,"depth":588,"text":5988},{"id":1910,"depth":587,"text":1913},"content:blog:llm-analytics-benchmark-leaderboard.md","blog/llm-analytics-benchmark-leaderboard.md","blog/llm-analytics-benchmark-leaderboard",{"_path":1329,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":6077,"description":6078,"draft":7,"publicationDate":915,"updatedAt":915,"image":6079,"author":6080,"ogTitle":6081,"ogDescription":6082,"twitterTitle":6083,"twitterDescription":6084,"keywords":6085,"tags":6086,"head":6091,"category":42,"body":6098,"_type":596,"_id":8053,"_source":598,"_file":8054,"_stem":8055,"_extension":601},"6 New LLMs, 3 Runs Each: The Best AI for Analytics Costs $0.06","We tested 6 new AI models 3 times each on the same marketing attribution query. MiniMax M2.5 delivered excellent results at $0.06 total, outperforming Claude Opus 4.6 at $1.35. One model couldn't even start.","/images/blog/llm-benchmark-round2-hero.svg",{"id":14,"name":15,"role":16,"twitter":17},"LLM Analytics Benchmark Round 2: 6 Models, 3 Runs Each","MiniMax M2.5 at $0.06 beat Claude Opus 4.6 at $1.35 across 3 consistency runs. Aurora Alpha failed all 3 attempts. See the full Round 2 benchmark results.","🔬 Round 2: 6 New LLMs Benchmarked 3x Each on Analytics","The best AI for analytics costs $0.06. MiniMax M2.5 dominated speed, cost, AND quality. Claude Opus 4.6 was the most thorough but 22x more expensive.","LLM benchmark round 2, AI analytics consistency test, MiniMax M2.5 review, Claude Opus 4.6 benchmark, Kimi K2.5 analytics, GLM 5 review, Qwen3 Max Thinking, Aurora Alpha, best LLM for analytics, AI marketing attribution, LLM cost comparison, AI consistency testing",[3209,25,930,933,6087,6088,6089,928,3213,6090],"Kimi","Qwen","GLM","consistency testing",{"meta":6092},[6093,6095,6096,6097],{"name":33,"content":6094},"LLM benchmark round 2, AI analytics consistency test, MiniMax M2.5, Claude Opus 4.6, Kimi K2.5, GLM 5, Qwen3 Max Thinking, Aurora Alpha, best AI for analytics, LLM consistency, marketing attribution AI",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},{"type":44,"children":6099,"toc":8021},[6100,6106,6160,6176,6181,6186,6209,6220,6226,6230,6235,6241,6614,6623,6629,6634,6831,6841,6847,6853,6859,6889,6894,6905,6912,6940,6950,6960,6966,6992,6997,7008,7016,7039,7048,7054,7080,7092,7103,7111,7139,7148,7158,7164,7190,7195,7206,7213,7236,7245,7251,7278,7283,7294,7301,7319,7345,7351,7357,7379,7384,7395,7405,7415,7421,7427,7572,7578,7707,7717,7722,7728,7732,7739,7768,7775,7808,7815,7848,7858,7861,7890,7894,7900,7911,7917,7922,7928,7933,7939,7944,7950,7955,7959,7964,7968,7971,7975],{"type":47,"tag":55,"props":6101,"children":6103},{"id":6102},"why-we-ran-each-model-3-times",[6104],{"type":52,"value":6105},"Why We Ran Each Model 3 Times",{"type":47,"tag":965,"props":6107,"children":6108},{"icon":967,"title":3231,"type":890},[6109],{"type":47,"tag":137,"props":6110,"children":6111},{},[6112,6121,6130,6140,6150],{"type":47,"tag":141,"props":6113,"children":6114},{},[6115,6119],{"type":47,"tag":88,"props":6116,"children":6117},{},[6118],{"type":52,"value":5253},{"type":52,"value":6120}," delivered excellent quality across all 3 runs",{"type":47,"tag":141,"props":6122,"children":6123},{},[6124,6128],{"type":47,"tag":88,"props":6125,"children":6126},{},[6127],{"type":52,"value":3263},{"type":52,"value":6129}," MiniMax M2.5 -- fastest (70s avg), cheapest ($0.06), excellent quality",{"type":47,"tag":141,"props":6131,"children":6132},{},[6133,6138],{"type":47,"tag":88,"props":6134,"children":6135},{},[6136],{"type":52,"value":6137},"Most thorough:",{"type":52,"value":6139}," Claude Opus 4.6 ($1.35) -- comprehensive analysis but 22x the cost",{"type":47,"tag":141,"props":6141,"children":6142},{},[6143,6148],{"type":47,"tag":88,"props":6144,"children":6145},{},[6146],{"type":52,"value":6147},"Complete failure:",{"type":52,"value":6149}," Aurora Alpha (stealth OpenRouter release) couldn't start -- context window too small",{"type":47,"tag":141,"props":6151,"children":6152},{},[6153,6158],{"type":47,"tag":88,"props":6154,"children":6155},{},[6156],{"type":52,"value":6157},"Consistency matters:",{"type":52,"value":6159}," Run times varied up to 90% within the same model across runs",{"type":47,"tag":48,"props":6161,"children":6162},{},[6163,6165,6169,6171],{"type":52,"value":6164},"In ",{"type":47,"tag":67,"props":6166,"children":6167},{"href":1313},[6168],{"type":52,"value":1316},{"type":52,"value":6170},", we tested 10 LLMs on broken GA4 data with a single run each. The results were revealing, but they left an important question unanswered: ",{"type":47,"tag":88,"props":6172,"children":6173},{},[6174],{"type":52,"value":6175},"How consistent are these models?",{"type":47,"tag":48,"props":6177,"children":6178},{},[6179],{"type":52,"value":6180},"LLMs are probabilistic systems. The same prompt can produce different results every time. A model that gives a brilliant answer once might stumble the next time. For analytics products, this is a critical concern. You can't ship an AI assistant that's excellent 60% of the time and mediocre the rest.",{"type":47,"tag":48,"props":6182,"children":6183},{},[6184],{"type":52,"value":6185},"So for Round 2, we changed two things:",{"type":47,"tag":522,"props":6187,"children":6188},{},[6189,6199],{"type":47,"tag":141,"props":6190,"children":6191},{},[6192,6197],{"type":47,"tag":88,"props":6193,"children":6194},{},[6195],{"type":52,"value":6196},"3 runs per model",{"type":52,"value":6198}," -- every model ran the same query 3 separate times",{"type":47,"tag":141,"props":6200,"children":6201},{},[6202,6207],{"type":47,"tag":88,"props":6203,"children":6204},{},[6205],{"type":52,"value":6206},"6 new models",{"type":52,"value":6208}," -- MiniMax M2.5, Kimi K2.5, Claude Opus 4.6, GLM 5, Qwen3 Max Thinking, and Aurora Alpha",{"type":47,"tag":48,"props":6210,"children":6211},{},[6212,6214,6218],{"type":52,"value":6213},"The test query remained the same: ",{"type":47,"tag":178,"props":6215,"children":6216},{},[6217],{"type":52,"value":3314},{"type":52,"value":6219}," And the same GA4 property with broken attribution data.",{"type":47,"tag":55,"props":6221,"children":6223},{"id":6222},"the-results-consistency-meets-cost-efficiency",[6224],{"type":52,"value":6225},"The Results: Consistency Meets Cost Efficiency",{"type":47,"tag":6227,"props":6228,"children":6229},"llm-benchmark-visualization-round2",{},[],{"type":47,"tag":1245,"props":6231,"children":6234},{"heading":6232,"icon":3338,"primary-link":1249,"primary-text":1250,"secondary-link":3339,"secondary-text":3340,"subheading":6233},"Want AI analytics you can rely on?","Anamap tests models rigorously so you don't have to. Get consistent, high-quality analytics insights.",[],{"type":47,"tag":55,"props":6236,"children":6238},{"id":6237},"full-leaderboard",[6239],{"type":52,"value":6240},"Full Leaderboard",{"type":47,"tag":1363,"props":6242,"children":6243},{},[6244,6295],{"type":47,"tag":1367,"props":6245,"children":6246},{},[6247],{"type":47,"tag":1371,"props":6248,"children":6249},{},[6250,6254,6258,6263,6267,6272,6277,6282,6286,6290],{"type":47,"tag":1375,"props":6251,"children":6252},{"align":3365},[6253],{"type":52,"value":3368},{"type":47,"tag":1375,"props":6255,"children":6256},{},[6257],{"type":52,"value":3373},{"type":47,"tag":1375,"props":6259,"children":6260},{"align":3365},[6261],{"type":52,"value":6262},"Runs",{"type":47,"tag":1375,"props":6264,"children":6265},{},[6266],{"type":52,"value":1629},{"type":47,"tag":1375,"props":6268,"children":6269},{"align":3365},[6270],{"type":52,"value":6271},"Accuracy",{"type":47,"tag":1375,"props":6273,"children":6274},{},[6275],{"type":52,"value":6276},"Avg Time",{"type":47,"tag":1375,"props":6278,"children":6279},{},[6280],{"type":52,"value":6281},"Best/Worst",{"type":47,"tag":1375,"props":6283,"children":6284},{},[6285],{"type":52,"value":3391},{"type":47,"tag":1375,"props":6287,"children":6288},{},[6289],{"type":52,"value":1485},{"type":47,"tag":1375,"props":6291,"children":6292},{},[6293],{"type":52,"value":6294},"$/1K tok",{"type":47,"tag":1391,"props":6296,"children":6297},{},[6298,6352,6404,6457,6510,6564],{"type":47,"tag":1371,"props":6299,"children":6300},{},[6301,6305,6313,6318,6322,6327,6332,6337,6342,6347],{"type":47,"tag":1398,"props":6302,"children":6303},{"align":3365},[6304],{"type":52,"value":3407},{"type":47,"tag":1398,"props":6306,"children":6307},{},[6308],{"type":47,"tag":67,"props":6309,"children":6311},{"href":6310},"#1-minimax-m25",[6312],{"type":52,"value":1070},{"type":47,"tag":1398,"props":6314,"children":6315},{"align":3365},[6316],{"type":52,"value":6317},"3/3",{"type":47,"tag":1398,"props":6319,"children":6320},{},[6321],{"type":52,"value":3421},{"type":47,"tag":1398,"props":6323,"children":6324},{"align":3365},[6325],{"type":52,"value":6326},"100",{"type":47,"tag":1398,"props":6328,"children":6329},{},[6330],{"type":52,"value":6331},"70s",{"type":47,"tag":1398,"props":6333,"children":6334},{},[6335],{"type":52,"value":6336},"52s / 88s",{"type":47,"tag":1398,"props":6338,"children":6339},{},[6340],{"type":52,"value":6341},"190K",{"type":47,"tag":1398,"props":6343,"children":6344},{},[6345],{"type":52,"value":6346},"$0.06",{"type":47,"tag":1398,"props":6348,"children":6349},{},[6350],{"type":52,"value":6351},"$0.0003",{"type":47,"tag":1371,"props":6353,"children":6354},{},[6355,6359,6367,6371,6375,6379,6384,6389,6394,6399],{"type":47,"tag":1398,"props":6356,"children":6357},{"align":3365},[6358],{"type":52,"value":3449},{"type":47,"tag":1398,"props":6360,"children":6361},{},[6362],{"type":47,"tag":67,"props":6363,"children":6365},{"href":6364},"#2-kimi-k25",[6366],{"type":52,"value":1226},{"type":47,"tag":1398,"props":6368,"children":6369},{"align":3365},[6370],{"type":52,"value":6317},{"type":47,"tag":1398,"props":6372,"children":6373},{},[6374],{"type":52,"value":3421},{"type":47,"tag":1398,"props":6376,"children":6377},{"align":3365},[6378],{"type":52,"value":6326},{"type":47,"tag":1398,"props":6380,"children":6381},{},[6382],{"type":52,"value":6383},"125s",{"type":47,"tag":1398,"props":6385,"children":6386},{},[6387],{"type":52,"value":6388},"108s / 145s",{"type":47,"tag":1398,"props":6390,"children":6391},{},[6392],{"type":52,"value":6393},"129K",{"type":47,"tag":1398,"props":6395,"children":6396},{},[6397],{"type":52,"value":6398},"$0.07",{"type":47,"tag":1398,"props":6400,"children":6401},{},[6402],{"type":52,"value":6403},"$0.0005",{"type":47,"tag":1371,"props":6405,"children":6406},{},[6407,6411,6420,6424,6428,6432,6437,6442,6447,6452],{"type":47,"tag":1398,"props":6408,"children":6409},{"align":3365},[6410],{"type":52,"value":3490},{"type":47,"tag":1398,"props":6412,"children":6413},{},[6414],{"type":47,"tag":67,"props":6415,"children":6417},{"href":6416},"#3-claude-opus-46",[6418],{"type":52,"value":6419},"Claude Opus 4.6",{"type":47,"tag":1398,"props":6421,"children":6422},{"align":3365},[6423],{"type":52,"value":6317},{"type":47,"tag":1398,"props":6425,"children":6426},{},[6427],{"type":52,"value":3421},{"type":47,"tag":1398,"props":6429,"children":6430},{"align":3365},[6431],{"type":52,"value":6326},{"type":47,"tag":1398,"props":6433,"children":6434},{},[6435],{"type":52,"value":6436},"143s",{"type":47,"tag":1398,"props":6438,"children":6439},{},[6440],{"type":52,"value":6441},"113s / 167s",{"type":47,"tag":1398,"props":6443,"children":6444},{},[6445],{"type":52,"value":6446},"238K",{"type":47,"tag":1398,"props":6448,"children":6449},{},[6450],{"type":52,"value":6451},"$1.35",{"type":47,"tag":1398,"props":6453,"children":6454},{},[6455],{"type":52,"value":6456},"$0.0056",{"type":47,"tag":1371,"props":6458,"children":6459},{},[6460,6464,6473,6477,6481,6485,6490,6495,6500,6505],{"type":47,"tag":1398,"props":6461,"children":6462},{"align":3365},[6463],{"type":52,"value":3530},{"type":47,"tag":1398,"props":6465,"children":6466},{},[6467],{"type":47,"tag":67,"props":6468,"children":6470},{"href":6469},"#4-glm-5",[6471],{"type":52,"value":6472},"GLM 5",{"type":47,"tag":1398,"props":6474,"children":6475},{"align":3365},[6476],{"type":52,"value":6317},{"type":47,"tag":1398,"props":6478,"children":6479},{},[6480],{"type":52,"value":3421},{"type":47,"tag":1398,"props":6482,"children":6483},{"align":3365},[6484],{"type":52,"value":6326},{"type":47,"tag":1398,"props":6486,"children":6487},{},[6488],{"type":52,"value":6489},"205s",{"type":47,"tag":1398,"props":6491,"children":6492},{},[6493],{"type":52,"value":6494},"145s / 275s",{"type":47,"tag":1398,"props":6496,"children":6497},{},[6498],{"type":52,"value":6499},"184K",{"type":47,"tag":1398,"props":6501,"children":6502},{},[6503],{"type":52,"value":6504},"$0.16",{"type":47,"tag":1398,"props":6506,"children":6507},{},[6508],{"type":52,"value":6509},"$0.0009",{"type":47,"tag":1371,"props":6511,"children":6512},{},[6513,6517,6526,6530,6534,6539,6544,6549,6554,6559],{"type":47,"tag":1398,"props":6514,"children":6515},{"align":3365},[6516],{"type":52,"value":3570},{"type":47,"tag":1398,"props":6518,"children":6519},{},[6520],{"type":47,"tag":67,"props":6521,"children":6523},{"href":6522},"#5-qwen3-max-thinking",[6524],{"type":52,"value":6525},"Qwen3 Max Thinking",{"type":47,"tag":1398,"props":6527,"children":6528},{"align":3365},[6529],{"type":52,"value":6317},{"type":47,"tag":1398,"props":6531,"children":6532},{},[6533],{"type":52,"value":3421},{"type":47,"tag":1398,"props":6535,"children":6536},{"align":3365},[6537],{"type":52,"value":6538},"96",{"type":47,"tag":1398,"props":6540,"children":6541},{},[6542],{"type":52,"value":6543},"89s",{"type":47,"tag":1398,"props":6545,"children":6546},{},[6547],{"type":52,"value":6548},"75s / 114s",{"type":47,"tag":1398,"props":6550,"children":6551},{},[6552],{"type":52,"value":6553},"359K",{"type":47,"tag":1398,"props":6555,"children":6556},{},[6557],{"type":52,"value":6558},"$0.44",{"type":47,"tag":1398,"props":6560,"children":6561},{},[6562],{"type":52,"value":6563},"$0.0012",{"type":47,"tag":1371,"props":6565,"children":6566},{},[6567,6572,6580,6585,6590,6594,6598,6602,6606,6610],{"type":47,"tag":1398,"props":6568,"children":6569},{"align":3365},[6570],{"type":52,"value":6571},"-",{"type":47,"tag":1398,"props":6573,"children":6574},{},[6575],{"type":47,"tag":67,"props":6576,"children":6578},{"href":6577},"#aurora-alpha-failed",[6579],{"type":52,"value":5263},{"type":47,"tag":1398,"props":6581,"children":6582},{"align":3365},[6583],{"type":52,"value":6584},"0/3",{"type":47,"tag":1398,"props":6586,"children":6587},{},[6588],{"type":52,"value":6589},"💥 error",{"type":47,"tag":1398,"props":6591,"children":6592},{"align":3365},[6593],{"type":52,"value":6571},{"type":47,"tag":1398,"props":6595,"children":6596},{},[6597],{"type":52,"value":6571},{"type":47,"tag":1398,"props":6599,"children":6600},{},[6601],{"type":52,"value":6571},{"type":47,"tag":1398,"props":6603,"children":6604},{},[6605],{"type":52,"value":6571},{"type":47,"tag":1398,"props":6607,"children":6608},{},[6609],{"type":52,"value":6571},{"type":47,"tag":1398,"props":6611,"children":6612},{},[6613],{"type":52,"value":6571},{"type":47,"tag":48,"props":6615,"children":6616},{},[6617,6621],{"type":47,"tag":88,"props":6618,"children":6619},{},[6620],{"type":52,"value":3812},{"type":52,"value":6622}," All 5 successful models achieved \"excellent\" quality ratings and near-perfect accuracy scores (96-100). The ranking is based on the combination of cost efficiency, speed, and analytical depth.",{"type":47,"tag":55,"props":6624,"children":6626},{"id":6625},"consistency-spotlight",[6627],{"type":52,"value":6628},"Consistency Spotlight",{"type":47,"tag":48,"props":6630,"children":6631},{},[6632],{"type":52,"value":6633},"One of the most interesting findings from running 3 tests per model was the variance in execution time, even when quality remained consistent.",{"type":47,"tag":1363,"props":6635,"children":6636},{},[6637,6672],{"type":47,"tag":1367,"props":6638,"children":6639},{},[6640],{"type":47,"tag":1371,"props":6641,"children":6642},{},[6643,6647,6652,6657,6662,6667],{"type":47,"tag":1375,"props":6644,"children":6645},{},[6646],{"type":52,"value":3373},{"type":47,"tag":1375,"props":6648,"children":6649},{},[6650],{"type":52,"value":6651},"Run 1",{"type":47,"tag":1375,"props":6653,"children":6654},{},[6655],{"type":52,"value":6656},"Run 2",{"type":47,"tag":1375,"props":6658,"children":6659},{},[6660],{"type":52,"value":6661},"Run 3",{"type":47,"tag":1375,"props":6663,"children":6664},{},[6665],{"type":52,"value":6666},"Variance",{"type":47,"tag":1375,"props":6668,"children":6669},{},[6670],{"type":52,"value":6671},"Quality Consistency",{"type":47,"tag":1391,"props":6673,"children":6674},{},[6675,6707,6738,6769,6800],{"type":47,"tag":1371,"props":6676,"children":6677},{},[6678,6682,6687,6692,6697,6702],{"type":47,"tag":1398,"props":6679,"children":6680},{},[6681],{"type":52,"value":1070},{"type":47,"tag":1398,"props":6683,"children":6684},{},[6685],{"type":52,"value":6686},"52s",{"type":47,"tag":1398,"props":6688,"children":6689},{},[6690],{"type":52,"value":6691},"69s",{"type":47,"tag":1398,"props":6693,"children":6694},{},[6695],{"type":52,"value":6696},"88s",{"type":47,"tag":1398,"props":6698,"children":6699},{},[6700],{"type":52,"value":6701},"69% spread",{"type":47,"tag":1398,"props":6703,"children":6704},{},[6705],{"type":52,"value":6706},"Excellent all 3",{"type":47,"tag":1371,"props":6708,"children":6709},{},[6710,6714,6719,6724,6729,6734],{"type":47,"tag":1398,"props":6711,"children":6712},{},[6713],{"type":52,"value":1226},{"type":47,"tag":1398,"props":6715,"children":6716},{},[6717],{"type":52,"value":6718},"108s",{"type":47,"tag":1398,"props":6720,"children":6721},{},[6722],{"type":52,"value":6723},"123s",{"type":47,"tag":1398,"props":6725,"children":6726},{},[6727],{"type":52,"value":6728},"145s",{"type":47,"tag":1398,"props":6730,"children":6731},{},[6732],{"type":52,"value":6733},"34% spread",{"type":47,"tag":1398,"props":6735,"children":6736},{},[6737],{"type":52,"value":6706},{"type":47,"tag":1371,"props":6739,"children":6740},{},[6741,6745,6750,6755,6760,6765],{"type":47,"tag":1398,"props":6742,"children":6743},{},[6744],{"type":52,"value":6419},{"type":47,"tag":1398,"props":6746,"children":6747},{},[6748],{"type":52,"value":6749},"113s",{"type":47,"tag":1398,"props":6751,"children":6752},{},[6753],{"type":52,"value":6754},"149s",{"type":47,"tag":1398,"props":6756,"children":6757},{},[6758],{"type":52,"value":6759},"167s",{"type":47,"tag":1398,"props":6761,"children":6762},{},[6763],{"type":52,"value":6764},"48% spread",{"type":47,"tag":1398,"props":6766,"children":6767},{},[6768],{"type":52,"value":6706},{"type":47,"tag":1371,"props":6770,"children":6771},{},[6772,6776,6781,6786,6791,6796],{"type":47,"tag":1398,"props":6773,"children":6774},{},[6775],{"type":52,"value":6525},{"type":47,"tag":1398,"props":6777,"children":6778},{},[6779],{"type":52,"value":6780},"75s",{"type":47,"tag":1398,"props":6782,"children":6783},{},[6784],{"type":52,"value":6785},"78s",{"type":47,"tag":1398,"props":6787,"children":6788},{},[6789],{"type":52,"value":6790},"114s",{"type":47,"tag":1398,"props":6792,"children":6793},{},[6794],{"type":52,"value":6795},"52% spread",{"type":47,"tag":1398,"props":6797,"children":6798},{},[6799],{"type":52,"value":6706},{"type":47,"tag":1371,"props":6801,"children":6802},{},[6803,6807,6811,6816,6821,6826],{"type":47,"tag":1398,"props":6804,"children":6805},{},[6806],{"type":52,"value":6472},{"type":47,"tag":1398,"props":6808,"children":6809},{},[6810],{"type":52,"value":6728},{"type":47,"tag":1398,"props":6812,"children":6813},{},[6814],{"type":52,"value":6815},"196s",{"type":47,"tag":1398,"props":6817,"children":6818},{},[6819],{"type":52,"value":6820},"275s",{"type":47,"tag":1398,"props":6822,"children":6823},{},[6824],{"type":52,"value":6825},"90% spread",{"type":47,"tag":1398,"props":6827,"children":6828},{},[6829],{"type":52,"value":6830},"2 excellent, 1 good",{"type":47,"tag":48,"props":6832,"children":6833},{},[6834,6839],{"type":47,"tag":88,"props":6835,"children":6836},{},[6837],{"type":52,"value":6838},"Key finding:",{"type":52,"value":6840}," Quality was remarkably stable across runs for most models. Even when execution time varied significantly (GLM 5 ranged from 145s to 275s), the analytical output remained consistent. The exception was GLM 5, where one run produced \"good\" rather than \"excellent\" quality -- the only quality variance observed across all 15 successful runs.",{"type":47,"tag":55,"props":6842,"children":6844},{"id":6843},"model-breakdowns",[6845],{"type":52,"value":6846},"Model Breakdowns",{"type":47,"tag":1057,"props":6848,"children":6850},{"id":6849},"models-that-delivered-excellent-results",[6851],{"type":52,"value":6852},"🏆 Models That Delivered Excellent Results",{"type":47,"tag":3840,"props":6854,"children":6856},{"id":6855},"_1-minimax-m25",[6857],{"type":52,"value":6858},"1. MiniMax M2.5",{"type":47,"tag":48,"props":6860,"children":6861},{},[6862,6866,6868,6873,6875,6880,6882,6887],{"type":47,"tag":88,"props":6863,"children":6864},{},[6865],{"type":52,"value":3853},{"type":52,"value":6867}," $0.06 | ",{"type":47,"tag":88,"props":6869,"children":6870},{},[6871],{"type":52,"value":6872},"Avg Time:",{"type":52,"value":6874}," 70s | ",{"type":47,"tag":88,"props":6876,"children":6877},{},[6878],{"type":52,"value":6879},"Accuracy:",{"type":52,"value":6881}," 100/100 | ",{"type":47,"tag":88,"props":6883,"children":6884},{},[6885],{"type":52,"value":6886},"Provider:",{"type":52,"value":6888}," MiniMax",{"type":47,"tag":48,"props":6890,"children":6891},{},[6892],{"type":52,"value":6893},"The runaway winner on every efficiency metric. MiniMax M2.5 was both the fastest and cheapest model while still delivering excellent analysis across all 3 runs.",{"type":47,"tag":755,"props":6895,"children":6896},{},[6897],{"type":47,"tag":48,"props":6898,"children":6899},{},[6900],{"type":47,"tag":178,"props":6901,"children":6902},{},[6903],{"type":52,"value":6904},"\"GA4 property has limited source tracking configured (sessionSourceMedium shows '(not set)' for all traffic), making traditional traffic source analysis unreliable. However, valuable conversion signals exist: subscription_upgrade (69,121 events), add_on_purchased (65,561), and sign_up (1,309).\"",{"type":47,"tag":48,"props":6906,"children":6907},{},[6908],{"type":47,"tag":88,"props":6909,"children":6910},{},[6911],{"type":52,"value":3886},{"type":47,"tag":137,"props":6913,"children":6914},{},[6915,6920,6925,6930,6935],{"type":47,"tag":141,"props":6916,"children":6917},{},[6918],{"type":52,"value":6919},"Immediately identified the attribution tracking gap",{"type":47,"tag":141,"props":6921,"children":6922},{},[6923],{"type":52,"value":6924},"Pivoted to conversion event analysis as the key value driver",{"type":47,"tag":141,"props":6926,"children":6927},{},[6928],{"type":52,"value":6929},"Found the top landing pages: /dashboard (206K sessions), /features (86K), /docs (35K), /pricing (32K)",{"type":47,"tag":141,"props":6931,"children":6932},{},[6933],{"type":52,"value":6934},"Flagged that traffic source tracking needs improvement for proper ROI analysis",{"type":47,"tag":141,"props":6936,"children":6937},{},[6938],{"type":52,"value":6939},"Did all of this in under 70 seconds at $0.02 per run",{"type":47,"tag":48,"props":6941,"children":6942},{},[6943,6948],{"type":47,"tag":88,"props":6944,"children":6945},{},[6946],{"type":52,"value":6947},"Consistency across runs:",{"type":52,"value":6949}," Quality remained excellent across all 3 runs. Time ranged from 52s to 88s, but the analytical output was consistently thorough and well-structured.",{"type":47,"tag":48,"props":6951,"children":6952},{},[6953,6958],{"type":47,"tag":88,"props":6954,"children":6955},{},[6956],{"type":52,"value":6957},"Cost comparison:",{"type":52,"value":6959}," At $0.0003 per 1K tokens, MiniMax M2.5 is 19x cheaper than Claude Opus 4.6 ($0.0056/1K). For teams running hundreds of queries per day, this difference is enormous.",{"type":47,"tag":3840,"props":6961,"children":6963},{"id":6962},"_2-kimi-k25",[6964],{"type":52,"value":6965},"2. Kimi K2.5",{"type":47,"tag":48,"props":6967,"children":6968},{},[6969,6973,6975,6979,6981,6985,6986,6990],{"type":47,"tag":88,"props":6970,"children":6971},{},[6972],{"type":52,"value":3853},{"type":52,"value":6974}," $0.07 | ",{"type":47,"tag":88,"props":6976,"children":6977},{},[6978],{"type":52,"value":6872},{"type":52,"value":6980}," 125s | ",{"type":47,"tag":88,"props":6982,"children":6983},{},[6984],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":6987,"children":6988},{},[6989],{"type":52,"value":6886},{"type":52,"value":6991}," MoonshotAI",{"type":47,"tag":48,"props":6993,"children":6994},{},[6995],{"type":52,"value":6996},"Kimi K2.5, from Chinese AI lab MoonshotAI, delivered a unique insight that no other model found.",{"type":47,"tag":755,"props":6998,"children":6999},{},[7000],{"type":47,"tag":48,"props":7001,"children":7002},{},[7003],{"type":47,"tag":178,"props":7004,"children":7005},{},[7006],{"type":52,"value":7007},"\"Landing page analysis reveals strong engagement on product marketing pages, with /landing/serverless achieving 98.5% engagement rate -- far outperforming the homepage.\"",{"type":47,"tag":48,"props":7009,"children":7010},{},[7011],{"type":47,"tag":88,"props":7012,"children":7013},{},[7014],{"type":52,"value":7015},"What stood out:",{"type":47,"tag":137,"props":7017,"children":7018},{},[7019,7024,7029,7034],{"type":47,"tag":141,"props":7020,"children":7021},{},[7022],{"type":52,"value":7023},"The only model to highlight the /landing/serverless page's exceptional engagement rate (98.5%)",{"type":47,"tag":141,"props":7025,"children":7026},{},[7027],{"type":52,"value":7028},"Noted that traffic source attribution data was unavailable, limiting full ROI analysis",{"type":47,"tag":141,"props":7030,"children":7031},{},[7032],{"type":52,"value":7033},"Maintained perfect accuracy scores across all runs",{"type":47,"tag":141,"props":7035,"children":7036},{},[7037],{"type":52,"value":7038},"Consistently solid at $0.07 total (just $0.01 more than MiniMax)",{"type":47,"tag":48,"props":7040,"children":7041},{},[7042,7046],{"type":47,"tag":88,"props":7043,"children":7044},{},[7045],{"type":52,"value":6947},{"type":52,"value":7047}," Most consistent timing of all models (108s-145s, 34% spread). Every run was excellent quality.",{"type":47,"tag":3840,"props":7049,"children":7051},{"id":7050},"_3-claude-opus-46",[7052],{"type":52,"value":7053},"3. Claude Opus 4.6",{"type":47,"tag":48,"props":7055,"children":7056},{},[7057,7061,7063,7067,7069,7073,7074,7078],{"type":47,"tag":88,"props":7058,"children":7059},{},[7060],{"type":52,"value":3853},{"type":52,"value":7062}," $1.35 | ",{"type":47,"tag":88,"props":7064,"children":7065},{},[7066],{"type":52,"value":6872},{"type":52,"value":7068}," 143s | ",{"type":47,"tag":88,"props":7070,"children":7071},{},[7072],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":7075,"children":7076},{},[7077],{"type":52,"value":6886},{"type":52,"value":7079}," Anthropic",{"type":47,"tag":48,"props":7081,"children":7082},{},[7083,7085,7090],{"type":52,"value":7084},"The most comprehensive analysis by far, but at 22x the cost of MiniMax M2.5. Claude Opus 4.6 (the newest version since our ",{"type":47,"tag":67,"props":7086,"children":7087},{"href":1313},[7088],{"type":52,"value":7089},"Round 1 test of Opus 4.5",{"type":52,"value":7091},") delivered the deepest investigation.",{"type":47,"tag":755,"props":7093,"children":7094},{},[7095],{"type":47,"tag":48,"props":7096,"children":7097},{},[7098],{"type":47,"tag":178,"props":7099,"children":7100},{},[7101],{"type":52,"value":7102},"\"The Dashboard (/dashboard) dominates session volume (979K sessions, 99.9% engagement rate), confirming strong product stickiness. For marketing-facing pages, /features leads with 1,004 active users and 428K sessions at 99.4% engagement.\"",{"type":47,"tag":48,"props":7104,"children":7105},{},[7106],{"type":47,"tag":88,"props":7107,"children":7108},{},[7109],{"type":52,"value":7110},"What made it thorough:",{"type":47,"tag":137,"props":7112,"children":7113},{},[7114,7119,7124,7129,7134],{"type":47,"tag":141,"props":7115,"children":7116},{},[7117],{"type":52,"value":7118},"Analyzed the entire conversion funnel: 6,337 sign_ups, 340,798 subscription_upgrades, 321,895 add-on purchases",{"type":47,"tag":141,"props":7120,"children":7121},{},[7122],{"type":52,"value":7123},"Quantified engagement rates per page: /pricing at 96.7%, /landing/serverless at 97.8%",{"type":47,"tag":141,"props":7125,"children":7126},{},[7127],{"type":52,"value":7128},"Explicitly called out the UTM tracking gap as a critical limitation",{"type":47,"tag":141,"props":7130,"children":7131},{},[7132],{"type":52,"value":7133},"Generated 4 data requests per run, investigating multiple angles",{"type":47,"tag":141,"props":7135,"children":7136},{},[7137],{"type":52,"value":7138},"Provided specific recommendations for fixing attribution tracking",{"type":47,"tag":48,"props":7140,"children":7141},{},[7142,7146],{"type":47,"tag":88,"props":7143,"children":7144},{},[7145],{"type":52,"value":6947},{"type":52,"value":7147}," Slightly wider time range (113s-167s) but quality was consistently excellent with near-perfect accuracy scores.",{"type":47,"tag":48,"props":7149,"children":7150},{},[7151,7156],{"type":47,"tag":88,"props":7152,"children":7153},{},[7154],{"type":52,"value":7155},"The cost question:",{"type":52,"value":7157}," Is Claude Opus 4.6's depth worth 22x the cost of MiniMax M2.5? For high-stakes strategy decisions, possibly. For routine daily queries, almost certainly not.",{"type":47,"tag":3840,"props":7159,"children":7161},{"id":7160},"_4-glm-5",[7162],{"type":52,"value":7163},"4. GLM 5",{"type":47,"tag":48,"props":7165,"children":7166},{},[7167,7171,7173,7177,7179,7183,7184,7188],{"type":47,"tag":88,"props":7168,"children":7169},{},[7170],{"type":52,"value":3853},{"type":52,"value":7172}," $0.16 | ",{"type":47,"tag":88,"props":7174,"children":7175},{},[7176],{"type":52,"value":6872},{"type":52,"value":7178}," 205s | ",{"type":47,"tag":88,"props":7180,"children":7181},{},[7182],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":7185,"children":7186},{},[7187],{"type":52,"value":6886},{"type":52,"value":7189}," Z.ai",{"type":47,"tag":48,"props":7191,"children":7192},{},[7193],{"type":52,"value":7194},"GLM 5 was the slowest model but provided the most actionable conversion rate analysis.",{"type":47,"tag":755,"props":7196,"children":7197},{},[7198],{"type":47,"tag":48,"props":7199,"children":7200},{},[7201],{"type":47,"tag":178,"props":7202,"children":7203},{},[7204],{"type":52,"value":7205},"\"Critical data quality issue discovered: all traffic source attribution is missing. However, landing page analysis reveals clear high-value drivers: /features page leads with 8,029 subscription upgrades (9.2% rate), /docs drives 4,943 upgrades (14.1% rate), and /pricing generates 3,317 upgrades (10.2% rate).\"",{"type":47,"tag":48,"props":7207,"children":7208},{},[7209],{"type":47,"tag":88,"props":7210,"children":7211},{},[7212],{"type":52,"value":4037},{"type":47,"tag":137,"props":7214,"children":7215},{},[7216,7221,7226,7231],{"type":47,"tag":141,"props":7217,"children":7218},{},[7219],{"type":52,"value":7220},"Calculated specific conversion rates per landing page -- the only model to do this",{"type":47,"tag":141,"props":7222,"children":7223},{},[7224],{"type":52,"value":7225},"Identified /docs at 14.1% upgrade rate as the highest-converting page",{"type":47,"tag":141,"props":7227,"children":7228},{},[7229],{"type":52,"value":7230},"Called out the UTM fix as \"immediate action required\"",{"type":47,"tag":141,"props":7232,"children":7233},{},[7234],{"type":52,"value":7235},"Provided a clear, actionable framework for marketing investment",{"type":47,"tag":48,"props":7237,"children":7238},{},[7239,7243],{"type":47,"tag":88,"props":7240,"children":7241},{},[7242],{"type":52,"value":6947},{"type":52,"value":7244}," This was the least consistent model. Run times ranged from 145s to 275s (90% spread), and one run scored \"good\" instead of \"excellent\" -- the only quality variance across all 15 successful runs in this benchmark.",{"type":47,"tag":3840,"props":7246,"children":7248},{"id":7247},"_5-qwen3-max-thinking",[7249],{"type":52,"value":7250},"5. Qwen3 Max Thinking",{"type":47,"tag":48,"props":7252,"children":7253},{},[7254,7258,7260,7264,7266,7270,7272,7276],{"type":47,"tag":88,"props":7255,"children":7256},{},[7257],{"type":52,"value":3853},{"type":52,"value":7259}," $0.44 | ",{"type":47,"tag":88,"props":7261,"children":7262},{},[7263],{"type":52,"value":6872},{"type":52,"value":7265}," 89s | ",{"type":47,"tag":88,"props":7267,"children":7268},{},[7269],{"type":52,"value":6879},{"type":52,"value":7271}," 96/100 | ",{"type":47,"tag":88,"props":7273,"children":7274},{},[7275],{"type":52,"value":6886},{"type":52,"value":7277}," Qwen",{"type":47,"tag":48,"props":7279,"children":7280},{},[7281],{"type":52,"value":7282},"Qwen3 Max Thinking was the second-fastest model and used the most tokens (359K), reflecting its \"thinking\" approach that processes internally before responding.",{"type":47,"tag":755,"props":7284,"children":7285},{},[7286],{"type":47,"tag":48,"props":7287,"children":7288},{},[7289],{"type":47,"tag":178,"props":7290,"children":7291},{},[7292],{"type":52,"value":7293},"\"The data shows that most traffic (100%) is coming from '(not set)' source, indicating a significant tracking issue with UTM parameters or referral data.\"",{"type":47,"tag":48,"props":7295,"children":7296},{},[7297],{"type":47,"tag":88,"props":7298,"children":7299},{},[7300],{"type":52,"value":7015},{"type":47,"tag":137,"props":7302,"children":7303},{},[7304,7309,7314],{"type":47,"tag":141,"props":7305,"children":7306},{},[7307],{"type":52,"value":7308},"Fast analysis despite high token usage (4,032 tokens/sec processing speed)",{"type":47,"tag":141,"props":7310,"children":7311},{},[7312],{"type":52,"value":7313},"4 data requests per run, matching Claude and MiniMax in investigation depth",{"type":47,"tag":141,"props":7315,"children":7316},{},[7317],{"type":52,"value":7318},"Correctly identified the core tracking issue",{"type":47,"tag":48,"props":7320,"children":7321},{},[7322,7327,7329,7335,7337,7343],{"type":47,"tag":88,"props":7323,"children":7324},{},[7325],{"type":52,"value":7326},"The accuracy caveat:",{"type":52,"value":7328}," Qwen3 scored 96/100 on accuracy rather than a perfect 100. It attempted to use 2 invalid GA4 dimensions (",{"type":47,"tag":332,"props":7330,"children":7332},{"className":7331},[],[7333],{"type":52,"value":7334},"audience_segment",{"type":52,"value":7336}," and ",{"type":47,"tag":332,"props":7338,"children":7340},{"className":7339},[],[7341],{"type":52,"value":7342},"audienceName",{"type":52,"value":7344},") that don't exist in the GA4 API. While these didn't derail the analysis, they indicate slightly less precise understanding of the GA4 schema compared to the top 4 models.",{"type":47,"tag":1057,"props":7346,"children":7348},{"id":7347},"failed-model",[7349],{"type":52,"value":7350},"❌ Failed Model",{"type":47,"tag":3840,"props":7352,"children":7354},{"id":7353},"aurora-alpha-failed",[7355],{"type":52,"value":7356},"Aurora Alpha (Failed)",{"type":47,"tag":48,"props":7358,"children":7359},{},[7360,7364,7366,7371,7373,7377],{"type":47,"tag":88,"props":7361,"children":7362},{},[7363],{"type":52,"value":3853},{"type":52,"value":7365}," N/A | ",{"type":47,"tag":88,"props":7367,"children":7368},{},[7369],{"type":52,"value":7370},"Runs:",{"type":52,"value":7372}," 0/3 | ",{"type":47,"tag":88,"props":7374,"children":7375},{},[7376],{"type":52,"value":6886},{"type":52,"value":7378}," Stealth release on OpenRouter (unknown backing provider)",{"type":47,"tag":48,"props":7380,"children":7381},{},[7382],{"type":52,"value":7383},"Aurora Alpha failed all 3 attempts before generating a single response. The system prompt (~126K tokens) plus the query exceeded its 128K context window limit.",{"type":47,"tag":755,"props":7385,"children":7386},{},[7387],{"type":47,"tag":48,"props":7388,"children":7389},{},[7390],{"type":47,"tag":178,"props":7391,"children":7392},{},[7393],{"type":52,"value":7394},"\"Context limit exceeded: 128000 tokens vs ~144K tokens required\"",{"type":47,"tag":48,"props":7396,"children":7397},{},[7398,7403],{"type":47,"tag":88,"props":7399,"children":7400},{},[7401],{"type":52,"value":7402},"Why this matters:",{"type":52,"value":7404}," Aurora Alpha appeared on OpenRouter as a stealth release with no publicly identified backing company and limited documentation. This lack of transparency made it impossible to verify its stated capabilities or troubleshoot the failure. The 128K context window is technically within range of the system prompt, but with zero headroom for conversation -- a fundamental limitation for any multi-turn analytics workflow.",{"type":47,"tag":48,"props":7406,"children":7407},{},[7408,7413],{"type":47,"tag":88,"props":7409,"children":7410},{},[7411],{"type":52,"value":7412},"Lesson for AI product builders:",{"type":52,"value":7414}," Always test context window limits under realistic conditions. A model that advertises 128K context but can't handle a typical analytics system prompt is not viable for production use, regardless of how well it performs on smaller prompts.",{"type":47,"tag":55,"props":7416,"children":7418},{"id":7417},"what-this-tells-us",[7419],{"type":52,"value":7420},"What This Tells Us",{"type":47,"tag":1057,"props":7422,"children":7424},{"id":7423},"the-cost-efficiency-revolution",[7425],{"type":52,"value":7426},"The Cost Efficiency Revolution",{"type":47,"tag":1363,"props":7428,"children":7429},{},[7430,7449],{"type":47,"tag":1367,"props":7431,"children":7432},{},[7433],{"type":47,"tag":1371,"props":7434,"children":7435},{},[7436,7440,7444],{"type":47,"tag":1375,"props":7437,"children":7438},{},[7439],{"type":52,"value":4406},{"type":47,"tag":1375,"props":7441,"children":7442},{},[7443],{"type":52,"value":4411},{"type":47,"tag":1375,"props":7445,"children":7446},{},[7447],{"type":52,"value":7448},"Runner-Up",{"type":47,"tag":1391,"props":7450,"children":7451},{},[7452,7472,7492,7512,7530,7551],{"type":47,"tag":1371,"props":7453,"children":7454},{},[7455,7462,7467],{"type":47,"tag":1398,"props":7456,"children":7457},{},[7458],{"type":47,"tag":88,"props":7459,"children":7460},{},[7461],{"type":52,"value":4427},{"type":47,"tag":1398,"props":7463,"children":7464},{},[7465],{"type":52,"value":7466},"MiniMax M2.5 (70s)",{"type":47,"tag":1398,"props":7468,"children":7469},{},[7470],{"type":52,"value":7471},"Qwen3 Max Thinking (89s)",{"type":47,"tag":1371,"props":7473,"children":7474},{},[7475,7482,7487],{"type":47,"tag":1398,"props":7476,"children":7477},{},[7478],{"type":47,"tag":88,"props":7479,"children":7480},{},[7481],{"type":52,"value":4445},{"type":47,"tag":1398,"props":7483,"children":7484},{},[7485],{"type":52,"value":7486},"MiniMax M2.5 ($0.06)",{"type":47,"tag":1398,"props":7488,"children":7489},{},[7490],{"type":52,"value":7491},"Kimi K2.5 ($0.07)",{"type":47,"tag":1371,"props":7493,"children":7494},{},[7495,7502,7507],{"type":47,"tag":1398,"props":7496,"children":7497},{},[7498],{"type":47,"tag":88,"props":7499,"children":7500},{},[7501],{"type":52,"value":4463},{"type":47,"tag":1398,"props":7503,"children":7504},{},[7505],{"type":52,"value":7506},"MiniMax M2.5 ($0.0003/1K)",{"type":47,"tag":1398,"props":7508,"children":7509},{},[7510],{"type":52,"value":7511},"Kimi K2.5 ($0.0005/1K)",{"type":47,"tag":1371,"props":7513,"children":7514},{},[7515,7522,7526],{"type":47,"tag":1398,"props":7516,"children":7517},{},[7518],{"type":47,"tag":88,"props":7519,"children":7520},{},[7521],{"type":52,"value":4498},{"type":47,"tag":1398,"props":7523,"children":7524},{},[7525],{"type":52,"value":6419},{"type":47,"tag":1398,"props":7527,"children":7528},{},[7529],{"type":52,"value":6472},{"type":47,"tag":1371,"props":7531,"children":7532},{},[7533,7541,7546],{"type":47,"tag":1398,"props":7534,"children":7535},{},[7536],{"type":47,"tag":88,"props":7537,"children":7538},{},[7539],{"type":52,"value":7540},"Most Consistent",{"type":47,"tag":1398,"props":7542,"children":7543},{},[7544],{"type":52,"value":7545},"Kimi K2.5 (34% spread)",{"type":47,"tag":1398,"props":7547,"children":7548},{},[7549],{"type":52,"value":7550},"Claude Opus 4.6 (48% spread)",{"type":47,"tag":1371,"props":7552,"children":7553},{},[7554,7562,7567],{"type":47,"tag":1398,"props":7555,"children":7556},{},[7557],{"type":47,"tag":88,"props":7558,"children":7559},{},[7560],{"type":52,"value":7561},"Unique Insight",{"type":47,"tag":1398,"props":7563,"children":7564},{},[7565],{"type":52,"value":7566},"Kimi K2.5 (98.5% engagement)",{"type":47,"tag":1398,"props":7568,"children":7569},{},[7570],{"type":52,"value":7571},"GLM 5 (conversion rates)",{"type":47,"tag":1057,"props":7573,"children":7575},{"id":7574},"round-1-vs-round-2",[7576],{"type":52,"value":7577},"Round 1 vs Round 2",{"type":47,"tag":1363,"props":7579,"children":7580},{},[7581,7600],{"type":47,"tag":1367,"props":7582,"children":7583},{},[7584],{"type":47,"tag":1371,"props":7585,"children":7586},{},[7587,7592,7596],{"type":47,"tag":1375,"props":7588,"children":7589},{},[7590],{"type":52,"value":7591},"Aspect",{"type":47,"tag":1375,"props":7593,"children":7594},{},[7595],{"type":52,"value":1316},{"type":47,"tag":1375,"props":7597,"children":7598},{},[7599],{"type":52,"value":1332},{"type":47,"tag":1391,"props":7601,"children":7602},{},[7603,7620,7637,7653,7671,7689],{"type":47,"tag":1371,"props":7604,"children":7605},{},[7606,7610,7615],{"type":47,"tag":1398,"props":7607,"children":7608},{},[7609],{"type":52,"value":1624},{"type":47,"tag":1398,"props":7611,"children":7612},{},[7613],{"type":52,"value":7614},"10 (established players)",{"type":47,"tag":1398,"props":7616,"children":7617},{},[7618],{"type":52,"value":7619},"6 (newer/niche models)",{"type":47,"tag":1371,"props":7621,"children":7622},{},[7623,7627,7632],{"type":47,"tag":1398,"props":7624,"children":7625},{},[7626],{"type":52,"value":6262},{"type":47,"tag":1398,"props":7628,"children":7629},{},[7630],{"type":52,"value":7631},"1 per model",{"type":47,"tag":1398,"props":7633,"children":7634},{},[7635],{"type":52,"value":7636},"3 per model",{"type":47,"tag":1371,"props":7638,"children":7639},{},[7640,7645,7649],{"type":47,"tag":1398,"props":7641,"children":7642},{},[7643],{"type":52,"value":7644},"Top performer",{"type":47,"tag":1398,"props":7646,"children":7647},{},[7648],{"type":52,"value":4455},{"type":47,"tag":1398,"props":7650,"children":7651},{},[7652],{"type":52,"value":7486},{"type":47,"tag":1371,"props":7654,"children":7655},{},[7656,7661,7666],{"type":47,"tag":1398,"props":7657,"children":7658},{},[7659],{"type":52,"value":7660},"Quality spread",{"type":47,"tag":1398,"props":7662,"children":7663},{},[7664],{"type":52,"value":7665},"30% excellent, 40% diagnostic, 30% hallucinated",{"type":47,"tag":1398,"props":7667,"children":7668},{},[7669],{"type":52,"value":7670},"83% excellent, 17% failed",{"type":47,"tag":1371,"props":7672,"children":7673},{},[7674,7679,7684],{"type":47,"tag":1398,"props":7675,"children":7676},{},[7677],{"type":52,"value":7678},"Accuracy range",{"type":47,"tag":1398,"props":7680,"children":7681},{},[7682],{"type":52,"value":7683},"75-100",{"type":47,"tag":1398,"props":7685,"children":7686},{},[7687],{"type":52,"value":7688},"96-100 (excluding failures)",{"type":47,"tag":1371,"props":7690,"children":7691},{},[7692,7697,7702],{"type":47,"tag":1398,"props":7693,"children":7694},{},[7695],{"type":52,"value":7696},"Key finding",{"type":47,"tag":1398,"props":7698,"children":7699},{},[7700],{"type":52,"value":7701},"Data quality judgment separates models",{"type":47,"tag":1398,"props":7703,"children":7704},{},[7705],{"type":52,"value":7706},"Cost efficiency doesn't sacrifice quality",{"type":47,"tag":48,"props":7708,"children":7709},{},[7710,7712],{"type":52,"value":7711},"The analytics AI market is maturing fast. In Round 1, model quality varied dramatically -- from fabricating data to providing genius-level analysis. In Round 2, quality has largely converged at the top. ",{"type":47,"tag":88,"props":7713,"children":7714},{},[7715],{"type":52,"value":7716},"The differentiation is now cost, speed, and consistency.",{"type":47,"tag":48,"props":7718,"children":7719},{},[7720],{"type":52,"value":7721},"A model that costs $0.06 and delivers excellent results 3 out of 3 times changes the economics of AI-powered analytics entirely.",{"type":47,"tag":1245,"props":7723,"children":7727},{"heading":7724,"icon":7725,"primary-link":1040,"primary-text":1251,"secondary-link":3339,"secondary-text":3340,"subheading":7726},"See the combined leaderboard","mdi-trophy-outline","16 models tested across 2 rounds. Find the right AI model for your analytics needs.",[],{"type":47,"tag":55,"props":7729,"children":7730},{"id":4603},[7731],{"type":52,"value":4606},{"type":47,"tag":48,"props":7733,"children":7734},{},[7735],{"type":47,"tag":88,"props":7736,"children":7737},{},[7738],{"type":52,"value":4614},{"type":47,"tag":137,"props":7740,"children":7741},{},[7742,7747,7751,7755,7759],{"type":47,"tag":141,"props":7743,"children":7744},{},[7745],{"type":52,"value":7746},"Same GA4 property as Round 1 with broken attribution tracking",{"type":47,"tag":141,"props":7748,"children":7749},{},[7750],{"type":52,"value":4627},{"type":47,"tag":141,"props":7752,"children":7753},{},[7754],{"type":52,"value":4632},{"type":47,"tag":141,"props":7756,"children":7757},{},[7758],{"type":52,"value":4637},{"type":47,"tag":141,"props":7760,"children":7761},{},[7762,7766],{"type":47,"tag":88,"props":7763,"children":7764},{},[7765],{"type":52,"value":6196},{"type":52,"value":7767}," to test consistency",{"type":47,"tag":48,"props":7769,"children":7770},{},[7771],{"type":47,"tag":88,"props":7772,"children":7773},{},[7774],{"type":52,"value":4645},{"type":47,"tag":137,"props":7776,"children":7777},{},[7778,7783,7788,7793,7798,7803],{"type":47,"tag":141,"props":7779,"children":7780},{},[7781],{"type":52,"value":7782},"MiniMax: MiniMax M2.5",{"type":47,"tag":141,"props":7784,"children":7785},{},[7786],{"type":52,"value":7787},"MoonshotAI: Kimi K2.5",{"type":47,"tag":141,"props":7789,"children":7790},{},[7791],{"type":52,"value":7792},"Anthropic: Claude Opus 4.6",{"type":47,"tag":141,"props":7794,"children":7795},{},[7796],{"type":52,"value":7797},"Z.ai: GLM 5",{"type":47,"tag":141,"props":7799,"children":7800},{},[7801],{"type":52,"value":7802},"Qwen: Qwen3 Max Thinking",{"type":47,"tag":141,"props":7804,"children":7805},{},[7806],{"type":52,"value":7807},"Aurora Alpha (stealth OpenRouter release, unknown provider)",{"type":47,"tag":48,"props":7809,"children":7810},{},[7811],{"type":47,"tag":88,"props":7812,"children":7813},{},[7814],{"type":52,"value":4681},{"type":47,"tag":137,"props":7816,"children":7817},{},[7818,7823,7828,7833,7838,7843],{"type":47,"tag":141,"props":7819,"children":7820},{},[7821],{"type":52,"value":7822},"Quality rating (excellent/good/fair/poor) across all 3 runs",{"type":47,"tag":141,"props":7824,"children":7825},{},[7826],{"type":52,"value":7827},"Accuracy score (GA4 field name accuracy, 0-100)",{"type":47,"tag":141,"props":7829,"children":7830},{},[7831],{"type":52,"value":7832},"Run time consistency (variance across 3 runs)",{"type":47,"tag":141,"props":7834,"children":7835},{},[7836],{"type":52,"value":7837},"Cost efficiency ($/1K tokens)",{"type":47,"tag":141,"props":7839,"children":7840},{},[7841],{"type":52,"value":7842},"Analytical depth and actionability",{"type":47,"tag":141,"props":7844,"children":7845},{},[7846],{"type":52,"value":7847},"Data quality issue detection and handling",{"type":47,"tag":48,"props":7849,"children":7850},{},[7851,7856],{"type":47,"tag":88,"props":7852,"children":7853},{},[7854],{"type":52,"value":7855},"Total benchmark cost:",{"type":52,"value":7857}," $2.08 across 18 runs (6 models x 3 runs)",{"type":47,"tag":1047,"props":7859,"children":7860},{},[],{"type":47,"tag":48,"props":7862,"children":7863},{},[7864],{"type":47,"tag":178,"props":7865,"children":7866},{},[7867,7869,7874,7876,7881,7883,7888],{"type":52,"value":7868},"This is Round 2 of our ongoing LLM analytics benchmark series conducted using the ",{"type":47,"tag":67,"props":7870,"children":7872},{"href":4723,"rel":7871},[71],[7873],{"type":52,"value":4727},{"type":52,"value":7875},". See ",{"type":47,"tag":67,"props":7877,"children":7878},{"href":1313},[7879],{"type":52,"value":7880},"Round 1: 10 Models on Broken Data",{"type":52,"value":7882}," for the original benchmark, or visit the ",{"type":47,"tag":67,"props":7884,"children":7885},{"href":1040},[7886],{"type":52,"value":7887},"combined leaderboard",{"type":52,"value":7889}," for all results across rounds.",{"type":47,"tag":55,"props":7891,"children":7892},{"id":1821},[7893],{"type":52,"value":1824},{"type":47,"tag":1057,"props":7895,"children":7897},{"id":7896},"which-ai-model-is-best-for-marketing-analytics-on-a-budget",[7898],{"type":52,"value":7899},"Which AI model is best for marketing analytics on a budget?",{"type":47,"tag":48,"props":7901,"children":7902},{},[7903,7905,7909],{"type":52,"value":7904},"Based on Round 2 results, ",{"type":47,"tag":88,"props":7906,"children":7907},{},[7908],{"type":52,"value":1070},{"type":52,"value":7910}," is the best budget option at $0.06 total across 3 runs ($0.02/query). It delivered excellent quality, perfect accuracy scores (100/100), and was the fastest model at 70 seconds average. At $0.0003 per 1K tokens, it's 19x cheaper than Claude Opus 4.6.",{"type":47,"tag":1057,"props":7912,"children":7914},{"id":7913},"how-consistent-are-llm-analytics-results",[7915],{"type":52,"value":7916},"How consistent are LLM analytics results?",{"type":47,"tag":48,"props":7918,"children":7919},{},[7920],{"type":52,"value":7921},"Remarkably consistent for quality, but variable for speed. In our 15 successful runs across 5 models, 14 out of 15 (93%) achieved \"excellent\" quality. However, execution time varied significantly -- GLM 5 ranged from 145s to 275s across its 3 runs. When choosing a model, expect the quality to be reliable but plan for timing variance.",{"type":47,"tag":1057,"props":7923,"children":7925},{"id":7924},"is-claude-opus-46-worth-the-premium-price",[7926],{"type":52,"value":7927},"Is Claude Opus 4.6 worth the premium price?",{"type":47,"tag":48,"props":7929,"children":7930},{},[7931],{"type":52,"value":7932},"It depends on the stakes. Claude Opus 4.6 ($1.35) delivered the most comprehensive analysis, investigating multiple angles with 4 data requests per run. But MiniMax M2.5 ($0.06) achieved the same \"excellent\" quality rating at 1/22nd the cost. For routine queries, the cheaper option wins. For high-stakes strategic decisions where depth matters, Claude's thoroughness justifies the premium.",{"type":47,"tag":1057,"props":7934,"children":7936},{"id":7935},"what-happened-to-aurora-alpha",[7937],{"type":52,"value":7938},"What happened to Aurora Alpha?",{"type":47,"tag":48,"props":7940,"children":7941},{},[7942],{"type":52,"value":7943},"Aurora Alpha is a stealth release on OpenRouter with no publicly identified backing company. It failed all 3 runs because its 128K context window couldn't accommodate the system prompt (~126K tokens) plus the query. This highlights the importance of testing real-world context requirements before adopting new models, especially those with limited public documentation.",{"type":47,"tag":1057,"props":7945,"children":7947},{"id":7946},"how-does-round-2-compare-to-round-1",[7948],{"type":52,"value":7949},"How does Round 2 compare to Round 1?",{"type":47,"tag":48,"props":7951,"children":7952},{},[7953],{"type":52,"value":7954},"Round 1 tested 10 established models (Claude, GPT-5, Gemini, Grok, DeepSeek) with a single run each. Quality varied dramatically: 30% delivered value, 40% stopped at diagnosis, 30% hallucinated. Round 2 tested 6 newer models 3 times each. Quality was much more consistent: 83% excellent. The key shift is that cost efficiency no longer means sacrificing quality.",{"type":47,"tag":1057,"props":7956,"children":7957},{"id":1894},[7958],{"type":52,"value":1897},{"type":47,"tag":48,"props":7960,"children":7961},{},[7962],{"type":52,"value":7963},"Three of our top 5 models were from Chinese AI labs: MiniMax M2.5 (#1), Kimi K2.5 (#2), and Qwen3 Max Thinking (#5). They delivered excellent quality at competitive prices. For analytics tasks that don't involve sensitive data, these models offer outstanding value. Consider data residency requirements if your organization has compliance constraints.",{"type":47,"tag":1904,"props":7965,"children":7967},{":faqs":7966},"[{\"question\":\"Which AI model is best for marketing analytics on a budget?\",\"answer\":\"MiniMax M2.5 is the best budget option at $0.06 total across 3 runs ($0.02/query). It delivered excellent quality, perfect accuracy scores (100/100), and was the fastest model at 70 seconds average. At $0.0003 per 1K tokens, it is 19x cheaper than Claude Opus 4.6.\"},{\"question\":\"How consistent are LLM analytics results?\",\"answer\":\"Remarkably consistent for quality, but variable for speed. In our 15 successful runs across 5 models, 14 out of 15 (93%) achieved excellent quality. However, execution time varied significantly -- GLM 5 ranged from 145s to 275s across its 3 runs.\"},{\"question\":\"Is Claude Opus 4.6 worth the premium price?\",\"answer\":\"It depends on the stakes. Claude Opus 4.6 ($1.35) delivered the most comprehensive analysis with 4 data requests per run. But MiniMax M2.5 ($0.06) achieved the same excellent quality rating at 1/22nd the cost. For routine queries, the cheaper option wins; for high-stakes decisions, Claude justifies the premium.\"},{\"question\":\"What happened to Aurora Alpha?\",\"answer\":\"Aurora Alpha is a stealth release on OpenRouter with no publicly identified backing company. It failed all 3 runs because its 128K context window could not accommodate the system prompt (~126K tokens) plus the query.\"},{\"question\":\"How does Round 2 compare to Round 1?\",\"answer\":\"Round 1 tested 10 established models with a single run each, and quality varied dramatically (30% hallucinated). Round 2 tested 6 newer models 3 times each, with 83% achieving excellent quality. The key shift is that cost efficiency no longer means sacrificing quality.\"},{\"question\":\"Should I use Chinese AI models for analytics?\",\"answer\":\"Three of our top 5 models were from Chinese AI labs: MiniMax M2.5 (#1), Kimi K2.5 (#2), and Qwen3 Max Thinking (#5). They delivered excellent quality at competitive prices. Consider data residency requirements if your organization has compliance constraints.\"}]",[],{"type":47,"tag":1047,"props":7969,"children":7970},{},[],{"type":47,"tag":55,"props":7972,"children":7973},{"id":1910},[7974],{"type":52,"value":1913},{"type":47,"tag":137,"props":7976,"children":7977},{},[7978,7986,7995,8004,8012],{"type":47,"tag":141,"props":7979,"children":7980},{},[7981,7985],{"type":47,"tag":67,"props":7982,"children":7983},{"href":912},[7984],{"type":52,"value":4926},{"type":52,"value":6016},{"type":47,"tag":141,"props":7987,"children":7988},{},[7989,7993],{"type":47,"tag":67,"props":7990,"children":7991},{"href":1313},[7992],{"type":52,"value":6032},{"type":52,"value":7994},": The original benchmark that started this series",{"type":47,"tag":141,"props":7996,"children":7997},{},[7998,8002],{"type":47,"tag":67,"props":7999,"children":8000},{"href":1040},[8001],{"type":52,"value":1924},{"type":52,"value":8003},": Combined results from all rounds in one place",{"type":47,"tag":141,"props":8005,"children":8006},{},[8007,8011],{"type":47,"tag":67,"props":8008,"children":8009},{"href":1962},[8010],{"type":52,"value":1965},{"type":52,"value":1967},{"type":47,"tag":141,"props":8013,"children":8014},{},[8015,8019],{"type":47,"tag":67,"props":8016,"children":8017},{"href":2200},[8018],{"type":52,"value":4964},{"type":52,"value":8020},": How to set up your analytics to avoid \"(not set)\" nightmares",{"title":8,"searchDepth":587,"depth":588,"links":8022},[8023,8024,8025,8026,8027,8039,8043,8044,8052],{"id":6102,"depth":587,"text":6105},{"id":6222,"depth":587,"text":6225},{"id":6237,"depth":587,"text":6240},{"id":6625,"depth":587,"text":6628},{"id":6843,"depth":587,"text":6846,"children":8028},[8029,8036],{"id":6849,"depth":588,"text":6852,"children":8030},[8031,8032,8033,8034,8035],{"id":6855,"depth":4986,"text":6858},{"id":6962,"depth":4986,"text":6965},{"id":7050,"depth":4986,"text":7053},{"id":7160,"depth":4986,"text":7163},{"id":7247,"depth":4986,"text":7250},{"id":7347,"depth":588,"text":7350,"children":8037},[8038],{"id":7353,"depth":4986,"text":7356},{"id":7417,"depth":587,"text":7420,"children":8040},[8041,8042],{"id":7423,"depth":588,"text":7426},{"id":7574,"depth":588,"text":7577},{"id":4603,"depth":587,"text":4606},{"id":1821,"depth":587,"text":1824,"children":8045},[8046,8047,8048,8049,8050,8051],{"id":7896,"depth":588,"text":7899},{"id":7913,"depth":588,"text":7916},{"id":7924,"depth":588,"text":7927},{"id":7935,"depth":588,"text":7938},{"id":7946,"depth":588,"text":7949},{"id":1894,"depth":588,"text":1897},{"id":1910,"depth":587,"text":1913},"content:blog:llm-analytics-benchmark-round-2-consistency.md","blog/llm-analytics-benchmark-round-2-consistency.md","blog/llm-analytics-benchmark-round-2-consistency",{"_path":1344,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":8057,"description":8058,"draft":7,"publicationDate":916,"updatedAt":916,"image":917,"author":8059,"ogTitle":8060,"ogDescription":8061,"twitterTitle":8062,"twitterDescription":8063,"keywords":8064,"tags":8065,"head":8067,"category":42,"body":8074,"_type":596,"_id":9780,"_source":598,"_file":9781,"_stem":9782,"_extension":601},"Can AI Tell If Analytics Data Is Synthetic? 10 New LLMs Tested","We tested 10 newer frontier models on whether they could determine if a GA4 demo dataset was real or synthetic. Gemini 3.5 Flash gave the best evidence-backed audit, Grok 4.3 was fastest, and GPT-5.5 was useful but expensive.",{"id":14,"name":15,"role":16,"twitter":17},"LLM Analytics Benchmark Round 3: Real vs Synthetic Data","10 newer LLMs tested 3 times each on whether they could identify synthetic GA4 demo data and explain the evidence.","Can AI Tell If Analytics Data Is Synthetic?","We tested GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, MiniMax M3, Qwen3.7, GLM 5.2, and more.","LLM benchmark, AI analytics benchmark, synthetic data detection, GPT-5.5 benchmark, Claude Opus 4.8 benchmark, Gemini 3.5 Flash benchmark, Grok 4.3 benchmark, MiniMax M3, Qwen3.7, synthetic analytics data, GA4 AI",[3209,25,8066,931,930,932,3210,933,6088,928],"synthetic data",{"meta":8068},[8069,8071,8072,8073],{"name":33,"content":8070},"LLM benchmark, AI analytics benchmark, synthetic data detection, GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, MiniMax M3, Qwen3.7, GA4 AI",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},{"type":44,"children":8075,"toc":9750},[8076,8082,8149,8159,8169,8174,8182,8187,8192,8198,8203,8246,8258,8262,8842,8851,8858,8864,8869,8874,8912,8917,8921,8927,8953,8958,8963,8971,8994,9000,9025,9030,9035,9041,9065,9070,9075,9080,9086,9112,9117,9129,9135,9161,9166,9171,9177,9203,9208,9213,9219,9243,9248,9253,9259,9286,9291,9296,9316,9322,9347,9352,9357,9363,9387,9392,9397,9403,9408,9413,9419,9424,9429,9467,9472,9476,9484,9526,9536,9544,9586,9596,9602,9607,9634,9639,9644,9648,9654,9659,9665,9670,9676,9681,9687,9692,9698,9703],{"type":47,"tag":55,"props":8077,"children":8079},{"id":8078},"the-synthetic-data-test",[8080],{"type":52,"value":8081},"The Synthetic Data Test",{"type":47,"tag":965,"props":8083,"children":8085},{"icon":8084,"title":3231,"type":890},"mdi-flask-outline",[8086],{"type":47,"tag":137,"props":8087,"children":8088},{},[8089,8099,8109,8119,8129,8139],{"type":47,"tag":141,"props":8090,"children":8091},{},[8092,8097],{"type":47,"tag":88,"props":8093,"children":8094},{},[8095],{"type":52,"value":8096},"9 of 10 models completed successfully",{"type":52,"value":8098}," after replacing Claude Fable 5, which was unavailable after a U.S. export-control directive, with GPT-5.5",{"type":47,"tag":141,"props":8100,"children":8101},{},[8102,8107],{"type":47,"tag":88,"props":8103,"children":8104},{},[8105],{"type":52,"value":8106},"Best evidence-backed audit:",{"type":52,"value":8108}," Gemini 3.5 Flash -- perfect GA4 field accuracy, actual data requests, clear synthetic-data reasoning",{"type":47,"tag":141,"props":8110,"children":8111},{},[8112,8117],{"type":47,"tag":88,"props":8113,"children":8114},{},[8115],{"type":52,"value":8116},"Fastest classifier:",{"type":52,"value":8118}," Grok 4.3 -- excellent answer in 8 seconds, but it did not make data requests",{"type":47,"tag":141,"props":8120,"children":8121},{},[8122,8127],{"type":47,"tag":88,"props":8123,"children":8124},{},[8125],{"type":52,"value":8126},"Best budget result:",{"type":52,"value":8128}," Gemini 3.1 Flash Lite -- $0.027 per run, excellent quality, but lower schema accuracy",{"type":47,"tag":141,"props":8130,"children":8131},{},[8132,8137],{"type":47,"tag":88,"props":8133,"children":8134},{},[8135],{"type":52,"value":8136},"Most expensive:",{"type":52,"value":8138}," Claude Opus 4.8 Fast ($1.64) and GPT-5.5 ($1.45) were useful but not cost leaders",{"type":47,"tag":141,"props":8140,"children":8141},{},[8142,8147],{"type":47,"tag":88,"props":8143,"children":8144},{},[8145],{"type":52,"value":8146},"Main finding:",{"type":52,"value":8148}," Most models correctly identified the dataset as synthetic, but they varied a lot in how much evidence they actually gathered",{"type":47,"tag":48,"props":8150,"children":8151},{},[8152,8153,8157],{"type":52,"value":6164},{"type":47,"tag":67,"props":8154,"children":8155},{"href":1313},[8156],{"type":52,"value":1316},{"type":52,"value":8158},", we tested whether LLMs could avoid hallucinating when attribution data was broken.",{"type":47,"tag":48,"props":8160,"children":8161},{},[8162,8163,8167],{"type":52,"value":6164},{"type":47,"tag":67,"props":8164,"children":8165},{"href":1329},[8166],{"type":52,"value":1332},{"type":52,"value":8168},", we tested whether newer models could repeat a good analytics answer three times in a row.",{"type":47,"tag":48,"props":8170,"children":8171},{},[8172],{"type":52,"value":8173},"For Round 3, we changed the question:",{"type":47,"tag":755,"props":8175,"children":8176},{},[8177],{"type":47,"tag":48,"props":8178,"children":8179},{},[8180],{"type":52,"value":8181},"Determine whether the connected analytics dataset appears to be real production data, synthetic demo data, or inconclusive. Use only evidence available from the connected data sources.",{"type":47,"tag":48,"props":8183,"children":8184},{},[8185],{"type":52,"value":8186},"That is a different kind of analytics task. It is less about campaign recommendations and more about evidence quality. A good model needs to inspect the data, notice statistical and structural tells, avoid overclaiming, and explain what evidence would change its mind.",{"type":47,"tag":48,"props":8188,"children":8189},{},[8190],{"type":52,"value":8191},"This matters because demo data is not just a sales prop. Good synthetic data should help people understand a product, test workflows, write content, and catch model weaknesses. Bad synthetic data teaches models the wrong lesson.",{"type":47,"tag":55,"props":8193,"children":8195},{"id":8194},"the-prompt",[8196],{"type":52,"value":8197},"The Prompt",{"type":47,"tag":48,"props":8199,"children":8200},{},[8201],{"type":52,"value":8202},"We asked each model to return:",{"type":47,"tag":522,"props":8204,"children":8205},{},[8206,8211,8216,8221,8226,8231,8236,8241],{"type":47,"tag":141,"props":8207,"children":8208},{},[8209],{"type":52,"value":8210},"Conclusion: real / synthetic / inconclusive",{"type":47,"tag":141,"props":8212,"children":8213},{},[8214],{"type":52,"value":8215},"Confidence: 0-100",{"type":47,"tag":141,"props":8217,"children":8218},{},[8219],{"type":52,"value":8220},"Evidence supporting the conclusion",{"type":47,"tag":141,"props":8222,"children":8223},{},[8224],{"type":52,"value":8225},"Evidence against the conclusion",{"type":47,"tag":141,"props":8227,"children":8228},{},[8229],{"type":52,"value":8230},"Specific synthetic-data tells, if any",{"type":47,"tag":141,"props":8232,"children":8233},{},[8234],{"type":52,"value":8235},"Specific real-data tells, if any",{"type":47,"tag":141,"props":8237,"children":8238},{},[8239],{"type":52,"value":8240},"What additional data would change its mind",{"type":47,"tag":141,"props":8242,"children":8243},{},[8244],{"type":52,"value":8245},"Recommendations to improve the demo data so it looks more realistic without becoming misleading",{"type":47,"tag":48,"props":8247,"children":8248},{},[8249,8251,8256],{"type":52,"value":8250},"Each model ran ",{"type":47,"tag":88,"props":8252,"children":8253},{},[8254],{"type":52,"value":8255},"3 times",{"type":52,"value":8257}," against the same connected GA4 property and the same Anamap analytics context.",{"type":47,"tag":55,"props":8259,"children":8260},{"id":6237},[8261],{"type":52,"value":6240},{"type":47,"tag":1363,"props":8263,"children":8264},{},[8265,8312],{"type":47,"tag":1367,"props":8266,"children":8267},{},[8268],{"type":47,"tag":1371,"props":8269,"children":8270},{},[8271,8275,8279,8283,8287,8291,8295,8299,8303,8307],{"type":47,"tag":1375,"props":8272,"children":8273},{"align":3365},[8274],{"type":52,"value":3368},{"type":47,"tag":1375,"props":8276,"children":8277},{},[8278],{"type":52,"value":3373},{"type":47,"tag":1375,"props":8280,"children":8281},{"align":3365},[8282],{"type":52,"value":6262},{"type":47,"tag":1375,"props":8284,"children":8285},{},[8286],{"type":52,"value":1629},{"type":47,"tag":1375,"props":8288,"children":8289},{"align":3365},[8290],{"type":52,"value":6271},{"type":47,"tag":1375,"props":8292,"children":8293},{},[8294],{"type":52,"value":6276},{"type":47,"tag":1375,"props":8296,"children":8297},{},[8298],{"type":52,"value":6281},{"type":47,"tag":1375,"props":8300,"children":8301},{},[8302],{"type":52,"value":3391},{"type":47,"tag":1375,"props":8304,"children":8305},{},[8306],{"type":52,"value":1485},{"type":47,"tag":1375,"props":8308,"children":8309},{},[8310],{"type":52,"value":8311},"Notes",{"type":47,"tag":1391,"props":8313,"children":8314},{},[8315,8367,8420,8471,8523,8577,8631,8682,8735,8789],{"type":47,"tag":1371,"props":8316,"children":8317},{},[8318,8322,8330,8334,8338,8342,8347,8352,8357,8362],{"type":47,"tag":1398,"props":8319,"children":8320},{"align":3365},[8321],{"type":52,"value":3407},{"type":47,"tag":1398,"props":8323,"children":8324},{},[8325],{"type":47,"tag":67,"props":8326,"children":8328},{"href":8327},"#1-gemini-35-flash",[8329],{"type":52,"value":1107},{"type":47,"tag":1398,"props":8331,"children":8332},{"align":3365},[8333],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8335,"children":8336},{},[8337],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8339,"children":8340},{"align":3365},[8341],{"type":52,"value":6326},{"type":47,"tag":1398,"props":8343,"children":8344},{},[8345],{"type":52,"value":8346},"53s",{"type":47,"tag":1398,"props":8348,"children":8349},{},[8350],{"type":52,"value":8351},"41s / 62s",{"type":47,"tag":1398,"props":8353,"children":8354},{},[8355],{"type":52,"value":8356},"98K",{"type":47,"tag":1398,"props":8358,"children":8359},{},[8360],{"type":52,"value":8361},"$0.23",{"type":47,"tag":1398,"props":8363,"children":8364},{},[8365],{"type":52,"value":8366},"Best evidence-backed audit",{"type":47,"tag":1371,"props":8368,"children":8369},{},[8370,8374,8383,8387,8391,8395,8400,8405,8410,8415],{"type":47,"tag":1398,"props":8371,"children":8372},{"align":3365},[8373],{"type":52,"value":3449},{"type":47,"tag":1398,"props":8375,"children":8376},{},[8377],{"type":47,"tag":67,"props":8378,"children":8380},{"href":8379},"#2-qwen37-max",[8381],{"type":52,"value":8382},"Qwen3.7 Max",{"type":47,"tag":1398,"props":8384,"children":8385},{"align":3365},[8386],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8388,"children":8389},{},[8390],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8392,"children":8393},{"align":3365},[8394],{"type":52,"value":6326},{"type":47,"tag":1398,"props":8396,"children":8397},{},[8398],{"type":52,"value":8399},"158s",{"type":47,"tag":1398,"props":8401,"children":8402},{},[8403],{"type":52,"value":8404},"128s / 198s",{"type":47,"tag":1398,"props":8406,"children":8407},{},[8408],{"type":52,"value":8409},"80K",{"type":47,"tag":1398,"props":8411,"children":8412},{},[8413],{"type":52,"value":8414},"$0.12",{"type":47,"tag":1398,"props":8416,"children":8417},{},[8418],{"type":52,"value":8419},"Detailed realism audit",{"type":47,"tag":1371,"props":8421,"children":8422},{},[8423,8427,8435,8439,8443,8447,8452,8457,8462,8466],{"type":47,"tag":1398,"props":8424,"children":8425},{"align":3365},[8426],{"type":52,"value":3490},{"type":47,"tag":1398,"props":8428,"children":8429},{},[8430],{"type":47,"tag":67,"props":8431,"children":8433},{"href":8432},"#3-minimax-m3",[8434],{"type":52,"value":1205},{"type":47,"tag":1398,"props":8436,"children":8437},{"align":3365},[8438],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8440,"children":8441},{},[8442],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8444,"children":8445},{"align":3365},[8446],{"type":52,"value":6326},{"type":47,"tag":1398,"props":8448,"children":8449},{},[8450],{"type":52,"value":8451},"250s",{"type":47,"tag":1398,"props":8453,"children":8454},{},[8455],{"type":52,"value":8456},"186s / 289s",{"type":47,"tag":1398,"props":8458,"children":8459},{},[8460],{"type":52,"value":8461},"141K",{"type":47,"tag":1398,"props":8463,"children":8464},{},[8465],{"type":52,"value":6346},{"type":47,"tag":1398,"props":8467,"children":8468},{},[8469],{"type":52,"value":8470},"Nuanced conclusion, slow",{"type":47,"tag":1371,"props":8472,"children":8473},{},[8474,8478,8486,8490,8494,8498,8503,8508,8513,8518],{"type":47,"tag":1398,"props":8475,"children":8476},{"align":3365},[8477],{"type":52,"value":3530},{"type":47,"tag":1398,"props":8479,"children":8480},{},[8481],{"type":47,"tag":67,"props":8482,"children":8484},{"href":8483},"#4-grok-43",[8485],{"type":52,"value":5168},{"type":47,"tag":1398,"props":8487,"children":8488},{"align":3365},[8489],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8491,"children":8492},{},[8493],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8495,"children":8496},{"align":3365},[8497],{"type":52,"value":6326},{"type":47,"tag":1398,"props":8499,"children":8500},{},[8501],{"type":52,"value":8502},"8s",{"type":47,"tag":1398,"props":8504,"children":8505},{},[8506],{"type":52,"value":8507},"5s / 10s",{"type":47,"tag":1398,"props":8509,"children":8510},{},[8511],{"type":52,"value":8512},"33K",{"type":47,"tag":1398,"props":8514,"children":8515},{},[8516],{"type":52,"value":8517},"$0.04",{"type":47,"tag":1398,"props":8519,"children":8520},{},[8521],{"type":52,"value":8522},"Fastest, but no data requests",{"type":47,"tag":1371,"props":8524,"children":8525},{},[8526,8530,8539,8543,8547,8552,8557,8562,8567,8572],{"type":47,"tag":1398,"props":8527,"children":8528},{"align":3365},[8529],{"type":52,"value":3570},{"type":47,"tag":1398,"props":8531,"children":8532},{},[8533],{"type":47,"tag":67,"props":8534,"children":8536},{"href":8535},"#5-claude-opus-48",[8537],{"type":52,"value":8538},"Claude Opus 4.8",{"type":47,"tag":1398,"props":8540,"children":8541},{"align":3365},[8542],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8544,"children":8545},{},[8546],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8548,"children":8549},{"align":3365},[8550],{"type":52,"value":8551},"94",{"type":47,"tag":1398,"props":8553,"children":8554},{},[8555],{"type":52,"value":8556},"81s",{"type":47,"tag":1398,"props":8558,"children":8559},{},[8560],{"type":52,"value":8561},"73s / 87s",{"type":47,"tag":1398,"props":8563,"children":8564},{},[8565],{"type":52,"value":8566},"135K",{"type":47,"tag":1398,"props":8568,"children":8569},{},[8570],{"type":52,"value":8571},"$0.81",{"type":47,"tag":1398,"props":8573,"children":8574},{},[8575],{"type":52,"value":8576},"Strong structural analysis",{"type":47,"tag":1371,"props":8578,"children":8579},{},[8580,8584,8593,8597,8601,8606,8611,8616,8621,8626],{"type":47,"tag":1398,"props":8581,"children":8582},{"align":3365},[8583],{"type":52,"value":3611},{"type":47,"tag":1398,"props":8585,"children":8586},{},[8587],{"type":47,"tag":67,"props":8588,"children":8590},{"href":8589},"#6-claude-opus-48-fast",[8591],{"type":52,"value":8592},"Claude Opus 4.8 Fast",{"type":47,"tag":1398,"props":8594,"children":8595},{"align":3365},[8596],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8598,"children":8599},{},[8600],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8602,"children":8603},{"align":3365},[8604],{"type":52,"value":8605},"93",{"type":47,"tag":1398,"props":8607,"children":8608},{},[8609],{"type":52,"value":8610},"32s",{"type":47,"tag":1398,"props":8612,"children":8613},{},[8614],{"type":52,"value":8615},"31s / 34s",{"type":47,"tag":1398,"props":8617,"children":8618},{},[8619],{"type":52,"value":8620},"137K",{"type":47,"tag":1398,"props":8622,"children":8623},{},[8624],{"type":52,"value":8625},"$1.64",{"type":47,"tag":1398,"props":8627,"children":8628},{},[8629],{"type":52,"value":8630},"Fast but expensive",{"type":47,"tag":1371,"props":8632,"children":8633},{},[8634,8638,8646,8650,8654,8659,8663,8668,8673,8677],{"type":47,"tag":1398,"props":8635,"children":8636},{"align":3365},[8637],{"type":52,"value":3651},{"type":47,"tag":1398,"props":8639,"children":8640},{},[8641],{"type":47,"tag":67,"props":8642,"children":8644},{"href":8643},"#7-gemini-31-flash-lite",[8645],{"type":52,"value":1195},{"type":47,"tag":1398,"props":8647,"children":8648},{"align":3365},[8649],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8651,"children":8652},{},[8653],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8655,"children":8656},{"align":3365},[8657],{"type":52,"value":8658},"90",{"type":47,"tag":1398,"props":8660,"children":8661},{},[8662],{"type":52,"value":8502},{"type":47,"tag":1398,"props":8664,"children":8665},{},[8666],{"type":52,"value":8667},"7s / 9s",{"type":47,"tag":1398,"props":8669,"children":8670},{},[8671],{"type":52,"value":8672},"101K",{"type":47,"tag":1398,"props":8674,"children":8675},{},[8676],{"type":52,"value":3512},{"type":47,"tag":1398,"props":8678,"children":8679},{},[8680],{"type":52,"value":8681},"Cheapest successful run",{"type":47,"tag":1371,"props":8683,"children":8684},{},[8685,8689,8697,8701,8705,8710,8715,8720,8725,8730],{"type":47,"tag":1398,"props":8686,"children":8687},{"align":3365},[8688],{"type":52,"value":3691},{"type":47,"tag":1398,"props":8690,"children":8691},{},[8692],{"type":47,"tag":67,"props":8693,"children":8695},{"href":8694},"#8-gpt-55",[8696],{"type":52,"value":5187},{"type":47,"tag":1398,"props":8698,"children":8699},{"align":3365},[8700],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8702,"children":8703},{},[8704],{"type":52,"value":3421},{"type":47,"tag":1398,"props":8706,"children":8707},{"align":3365},[8708],{"type":52,"value":8709},"88",{"type":47,"tag":1398,"props":8711,"children":8712},{},[8713],{"type":52,"value":8714},"148s",{"type":47,"tag":1398,"props":8716,"children":8717},{},[8718],{"type":52,"value":8719},"105s / 172s",{"type":47,"tag":1398,"props":8721,"children":8722},{},[8723],{"type":52,"value":8724},"224K",{"type":47,"tag":1398,"props":8726,"children":8727},{},[8728],{"type":52,"value":8729},"$1.45",{"type":47,"tag":1398,"props":8731,"children":8732},{},[8733],{"type":52,"value":8734},"Useful but schema issues",{"type":47,"tag":1371,"props":8736,"children":8737},{},[8738,8742,8751,8755,8760,8764,8769,8774,8779,8784],{"type":47,"tag":1398,"props":8739,"children":8740},{"align":3365},[8741],{"type":52,"value":3732},{"type":47,"tag":1398,"props":8743,"children":8744},{},[8745],{"type":47,"tag":67,"props":8746,"children":8748},{"href":8747},"#9-glm-52",[8749],{"type":52,"value":8750},"GLM 5.2",{"type":47,"tag":1398,"props":8752,"children":8753},{"align":3365},[8754],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8756,"children":8757},{},[8758],{"type":52,"value":8759},"✅ good",{"type":47,"tag":1398,"props":8761,"children":8762},{"align":3365},[8763],{"type":52,"value":8658},{"type":47,"tag":1398,"props":8765,"children":8766},{},[8767],{"type":52,"value":8768},"84s",{"type":47,"tag":1398,"props":8770,"children":8771},{},[8772],{"type":52,"value":8773},"47s / 133s",{"type":47,"tag":1398,"props":8775,"children":8776},{},[8777],{"type":52,"value":8778},"126K",{"type":47,"tag":1398,"props":8780,"children":8781},{},[8782],{"type":52,"value":8783},"$0.20",{"type":47,"tag":1398,"props":8785,"children":8786},{},[8787],{"type":52,"value":8788},"Correct but less complete",{"type":47,"tag":1371,"props":8790,"children":8791},{},[8792,8796,8805,8809,8814,8818,8823,8828,8833,8837],{"type":47,"tag":1398,"props":8793,"children":8794},{"align":3365},[8795],{"type":52,"value":3772},{"type":47,"tag":1398,"props":8797,"children":8798},{},[8799],{"type":47,"tag":67,"props":8800,"children":8802},{"href":8801},"#10-qwen37-plus",[8803],{"type":52,"value":8804},"Qwen3.7 Plus",{"type":47,"tag":1398,"props":8806,"children":8807},{"align":3365},[8808],{"type":52,"value":6317},{"type":47,"tag":1398,"props":8810,"children":8811},{},[8812],{"type":52,"value":8813},"⚠️ fair",{"type":47,"tag":1398,"props":8815,"children":8816},{"align":3365},[8817],{"type":52,"value":6326},{"type":47,"tag":1398,"props":8819,"children":8820},{},[8821],{"type":52,"value":8822},"146s",{"type":47,"tag":1398,"props":8824,"children":8825},{},[8826],{"type":52,"value":8827},"140s / 153s",{"type":47,"tag":1398,"props":8829,"children":8830},{},[8831],{"type":52,"value":8832},"149K",{"type":47,"tag":1398,"props":8834,"children":8835},{},[8836],{"type":52,"value":3714},{"type":47,"tag":1398,"props":8838,"children":8839},{},[8840],{"type":52,"value":8841},"Accurate syntax, thinner evidence",{"type":47,"tag":48,"props":8843,"children":8844},{},[8845,8849],{"type":47,"tag":88,"props":8846,"children":8847},{},[8848],{"type":52,"value":3812},{"type":52,"value":8850}," The auto-generated leaderboard ranked Grok 4.3 first because it was fast, cheap, successful, and had perfect field accuracy. For this article, I rank Gemini 3.5 Flash higher because the task explicitly required evidence from connected data, and Grok completed without making data requests.",{"type":47,"tag":1245,"props":8852,"children":8857},{"heading":8853,"icon":8854,"primary-link":1249,"primary-text":1250,"secondary-link":1040,"secondary-text":8855,"subheading":8856},"Want AI analytics with evidence?","mdi-database-search-outline","View Leaderboard","Anamap is built around data-source queries, citations, and model behavior testing -- not just fluent answers.",[],{"type":47,"tag":55,"props":8859,"children":8861},{"id":8860},"what-the-models-found",[8862],{"type":52,"value":8863},"What the Models Found",{"type":47,"tag":48,"props":8865,"children":8866},{},[8867],{"type":52,"value":8868},"The models mostly converged on the right answer: the dataset looks synthetic.",{"type":47,"tag":48,"props":8870,"children":8871},{},[8872],{"type":52,"value":8873},"The strongest evidence repeated across runs:",{"type":47,"tag":137,"props":8875,"children":8876},{},[8877,8882,8887,8892,8897,8902,8907],{"type":47,"tag":141,"props":8878,"children":8879},{},[8880],{"type":52,"value":8881},"The GA4 property is explicitly labeled as a test property.",{"type":47,"tag":141,"props":8883,"children":8884},{},[8885],{"type":52,"value":8886},"The requested historical window only returned a short populated date range.",{"type":47,"tag":141,"props":8888,"children":8889},{},[8890],{"type":52,"value":8891},"Traffic attribution is suspiciously incomplete or classified as unassigned.",{"type":47,"tag":141,"props":8893,"children":8894},{},[8895],{"type":52,"value":8896},"Geography, browser, and device distributions are too tidy.",{"type":47,"tag":141,"props":8898,"children":8899},{},[8900],{"type":52,"value":8901},"Event names and attributes match a curated tracking plan too closely.",{"type":47,"tag":141,"props":8903,"children":8904},{},[8905],{"type":52,"value":8906},"Conversion-like product events fire at high volume while GA4 conversions and revenue stay at zero.",{"type":47,"tag":141,"props":8908,"children":8909},{},[8910],{"type":52,"value":8911},"Some ratios are too smooth or too clean for production traffic.",{"type":47,"tag":48,"props":8913,"children":8914},{},[8915],{"type":52,"value":8916},"The best models did not stop at \"synthetic.\" They explained which signals were structural, which could be caused by broken tracking, and which additions would make the dataset more realistic.",{"type":47,"tag":55,"props":8918,"children":8919},{"id":6843},[8920],{"type":52,"value":6846},{"type":47,"tag":1057,"props":8922,"children":8924},{"id":8923},"_1-gemini-35-flash",[8925],{"type":52,"value":8926},"1. Gemini 3.5 Flash",{"type":47,"tag":48,"props":8928,"children":8929},{},[8930,8934,8936,8940,8942,8946,8947,8951],{"type":47,"tag":88,"props":8931,"children":8932},{},[8933],{"type":52,"value":3853},{"type":52,"value":8935}," $0.23 | ",{"type":47,"tag":88,"props":8937,"children":8938},{},[8939],{"type":52,"value":6872},{"type":52,"value":8941}," 53s | ",{"type":47,"tag":88,"props":8943,"children":8944},{},[8945],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":8948,"children":8949},{},[8950],{"type":52,"value":6886},{"type":52,"value":8952}," Google",{"type":47,"tag":48,"props":8954,"children":8955},{},[8956],{"type":52,"value":8957},"Gemini 3.5 Flash was the best balanced result for this specific task. It made data requests, kept perfect GA4 field accuracy, and clearly explained why the dataset was synthetic.",{"type":47,"tag":48,"props":8959,"children":8960},{},[8961],{"type":52,"value":8962},"Its final summary called out the lack of weekly seasonality, 100% missing traffic source attribution, mathematically suspicious event ratios, and alignment with static documentation benchmarks.",{"type":47,"tag":48,"props":8964,"children":8965},{},[8966],{"type":47,"tag":88,"props":8967,"children":8968},{},[8969],{"type":52,"value":8970},"Why it ranked first:",{"type":47,"tag":137,"props":8972,"children":8973},{},[8974,8979,8984,8989],{"type":47,"tag":141,"props":8975,"children":8976},{},[8977],{"type":52,"value":8978},"Perfect field accuracy",{"type":47,"tag":141,"props":8980,"children":8981},{},[8982],{"type":52,"value":8983},"Evidence-backed answer from queried data",{"type":47,"tag":141,"props":8985,"children":8986},{},[8987],{"type":52,"value":8988},"Good speed for a multi-turn audit",{"type":47,"tag":141,"props":8990,"children":8991},{},[8992],{"type":52,"value":8993},"Strong explanation of synthetic tells",{"type":47,"tag":1057,"props":8995,"children":8997},{"id":8996},"_2-qwen37-max",[8998],{"type":52,"value":8999},"2. Qwen3.7 Max",{"type":47,"tag":48,"props":9001,"children":9002},{},[9003,9007,9009,9013,9015,9019,9020,9024],{"type":47,"tag":88,"props":9004,"children":9005},{},[9006],{"type":52,"value":3853},{"type":52,"value":9008}," $0.12 | ",{"type":47,"tag":88,"props":9010,"children":9011},{},[9012],{"type":52,"value":6872},{"type":52,"value":9014}," 158s | ",{"type":47,"tag":88,"props":9016,"children":9017},{},[9018],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":9021,"children":9022},{},[9023],{"type":52,"value":6886},{"type":52,"value":7277},{"type":47,"tag":48,"props":9026,"children":9027},{},[9028],{"type":52,"value":9029},"Qwen3.7 Max gave one of the clearest realism audits. It identified exactly bounded geo values, limited browser diversity, a short temporal window, clean conversion funnels, and missing long-tail behavior.",{"type":47,"tag":48,"props":9031,"children":9032},{},[9033],{"type":52,"value":9034},"It was slower than Gemini 3.5 Flash, but its answer was easy to turn into a demo-data improvement checklist: add more geography, more browser/device tail, more temporal seasonality, and more realistic outliers.",{"type":47,"tag":1057,"props":9036,"children":9038},{"id":9037},"_3-minimax-m3",[9039],{"type":52,"value":9040},"3. MiniMax M3",{"type":47,"tag":48,"props":9042,"children":9043},{},[9044,9048,9049,9053,9055,9059,9060,9064],{"type":47,"tag":88,"props":9045,"children":9046},{},[9047],{"type":52,"value":3853},{"type":52,"value":6867},{"type":47,"tag":88,"props":9050,"children":9051},{},[9052],{"type":52,"value":6872},{"type":52,"value":9054}," 250s | ",{"type":47,"tag":88,"props":9056,"children":9057},{},[9058],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":9061,"children":9062},{},[9063],{"type":52,"value":6886},{"type":52,"value":6888},{"type":47,"tag":48,"props":9066,"children":9067},{},[9068],{"type":52,"value":9069},"MiniMax M3 was the most nuanced low-cost result. It concluded the dataset was likely synthetic, but it also noted counter-signals: work-hour skew, weekday dominance, event drift beyond the documented plan, realistic new/returning ratios, and plausible session-to-user ratios.",{"type":47,"tag":48,"props":9071,"children":9072},{},[9073],{"type":52,"value":9074},"That nuance is useful. A weaker model simply says \"synthetic\" and moves on. MiniMax described why the generator is already doing some things well.",{"type":47,"tag":48,"props":9076,"children":9077},{},[9078],{"type":52,"value":9079},"The tradeoff: it was very slow.",{"type":47,"tag":1057,"props":9081,"children":9083},{"id":9082},"_4-grok-43",[9084],{"type":52,"value":9085},"4. Grok 4.3",{"type":47,"tag":48,"props":9087,"children":9088},{},[9089,9093,9095,9099,9101,9105,9106,9110],{"type":47,"tag":88,"props":9090,"children":9091},{},[9092],{"type":52,"value":3853},{"type":52,"value":9094}," $0.04 | ",{"type":47,"tag":88,"props":9096,"children":9097},{},[9098],{"type":52,"value":6872},{"type":52,"value":9100}," 8s | ",{"type":47,"tag":88,"props":9102,"children":9103},{},[9104],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":9107,"children":9108},{},[9109],{"type":52,"value":6886},{"type":52,"value":9111}," xAI",{"type":47,"tag":48,"props":9113,"children":9114},{},[9115],{"type":52,"value":9116},"Grok 4.3 was the fastest successful model by a wide margin. It correctly classified the dataset as synthetic and gave a concise explanation.",{"type":47,"tag":48,"props":9118,"children":9119},{},[9120,9122,9127],{"type":52,"value":9121},"But the benchmark recorded ",{"type":47,"tag":88,"props":9123,"children":9124},{},[9125],{"type":52,"value":9126},"zero data requests",{"type":52,"value":9128},", which means it appears to have relied on provided context rather than actively inspecting the connected GA4 data. For a quick classifier, that is impressive. For an evidence audit, it is a limitation.",{"type":47,"tag":1057,"props":9130,"children":9132},{"id":9131},"_5-claude-opus-48",[9133],{"type":52,"value":9134},"5. Claude Opus 4.8",{"type":47,"tag":48,"props":9136,"children":9137},{},[9138,9142,9144,9148,9150,9154,9156,9160],{"type":47,"tag":88,"props":9139,"children":9140},{},[9141],{"type":52,"value":3853},{"type":52,"value":9143}," $0.81 | ",{"type":47,"tag":88,"props":9145,"children":9146},{},[9147],{"type":52,"value":6872},{"type":52,"value":9149}," 81s | ",{"type":47,"tag":88,"props":9151,"children":9152},{},[9153],{"type":52,"value":6879},{"type":52,"value":9155}," 94/100 | ",{"type":47,"tag":88,"props":9157,"children":9158},{},[9159],{"type":52,"value":6886},{"type":52,"value":7079},{"type":47,"tag":48,"props":9162,"children":9163},{},[9164],{"type":52,"value":9165},"Claude Opus 4.8 produced one of the strongest structural explanations. It identified geography collapsing to exact country/city pairs, browser distribution missing the normal long tail, a tidy cartesian grid across geo/device/browser, and zero new users across days.",{"type":47,"tag":48,"props":9167,"children":9168},{},[9169],{"type":52,"value":9170},"This was a rich answer, but it was not cheap. It also took a small accuracy hit from GA4 field issues.",{"type":47,"tag":1057,"props":9172,"children":9174},{"id":9173},"_6-claude-opus-48-fast",[9175],{"type":52,"value":9176},"6. Claude Opus 4.8 Fast",{"type":47,"tag":48,"props":9178,"children":9179},{},[9180,9184,9186,9190,9192,9196,9198,9202],{"type":47,"tag":88,"props":9181,"children":9182},{},[9183],{"type":52,"value":3853},{"type":52,"value":9185}," $1.64 | ",{"type":47,"tag":88,"props":9187,"children":9188},{},[9189],{"type":52,"value":6872},{"type":52,"value":9191}," 32s | ",{"type":47,"tag":88,"props":9193,"children":9194},{},[9195],{"type":52,"value":6879},{"type":52,"value":9197}," 93/100 | ",{"type":47,"tag":88,"props":9199,"children":9200},{},[9201],{"type":52,"value":6886},{"type":52,"value":7079},{"type":47,"tag":48,"props":9204,"children":9205},{},[9206],{"type":52,"value":9207},"Claude Opus 4.8 Fast lived up to the name on latency, but not on cost. It was fast and thoughtful, identifying smooth traffic, all-unassigned channels, a closed geography set, and suspicious product-event rates.",{"type":47,"tag":48,"props":9209,"children":9210},{},[9211],{"type":52,"value":9212},"The problem is economic: it cost more than GPT-5.5 and roughly 60x Gemini 3.1 Flash Lite.",{"type":47,"tag":1057,"props":9214,"children":9216},{"id":9215},"_7-gemini-31-flash-lite",[9217],{"type":52,"value":9218},"7. Gemini 3.1 Flash Lite",{"type":47,"tag":48,"props":9220,"children":9221},{},[9222,9226,9227,9231,9232,9236,9238,9242],{"type":47,"tag":88,"props":9223,"children":9224},{},[9225],{"type":52,"value":3853},{"type":52,"value":4007},{"type":47,"tag":88,"props":9228,"children":9229},{},[9230],{"type":52,"value":6872},{"type":52,"value":9100},{"type":47,"tag":88,"props":9233,"children":9234},{},[9235],{"type":52,"value":6879},{"type":52,"value":9237}," 90/100 | ",{"type":47,"tag":88,"props":9239,"children":9240},{},[9241],{"type":52,"value":6886},{"type":52,"value":8952},{"type":47,"tag":48,"props":9244,"children":9245},{},[9246],{"type":52,"value":9247},"Gemini 3.1 Flash Lite was the cheapest successful Round 3 model. It correctly identified synthetic/test-data patterns and was extremely fast.",{"type":47,"tag":48,"props":9249,"children":9250},{},[9251],{"type":52,"value":9252},"The answer was thinner than Gemini 3.5 Flash, and the field-accuracy score was lower. I would use it for cheap smoke tests, not as the final judge for a data-quality audit.",{"type":47,"tag":1057,"props":9254,"children":9256},{"id":9255},"_8-gpt-55",[9257],{"type":52,"value":9258},"8. GPT-5.5",{"type":47,"tag":48,"props":9260,"children":9261},{},[9262,9266,9268,9272,9274,9278,9280,9284],{"type":47,"tag":88,"props":9263,"children":9264},{},[9265],{"type":52,"value":3853},{"type":52,"value":9267}," $1.45 | ",{"type":47,"tag":88,"props":9269,"children":9270},{},[9271],{"type":52,"value":6872},{"type":52,"value":9273}," 148s | ",{"type":47,"tag":88,"props":9275,"children":9276},{},[9277],{"type":52,"value":6879},{"type":52,"value":9279}," 88/100 | ",{"type":47,"tag":88,"props":9281,"children":9282},{},[9283],{"type":52,"value":6886},{"type":52,"value":9285}," OpenAI",{"type":47,"tag":48,"props":9287,"children":9288},{},[9289],{"type":52,"value":9290},"GPT-5.5 was added as a one-off replacement after Claude Fable 5 failed all attempts through OpenRouter. Public reporting and Anthropic's own statement indicate Fable 5 had been disabled after a U.S. export-control directive, so we treated those failures as access-related rather than model-quality evidence.",{"type":47,"tag":48,"props":9292,"children":9293},{},[9294],{"type":52,"value":9295},"It correctly concluded the dataset was synthetic demo data. Its strongest evidence: the test-property label, only 8 populated days in a requested 90-day window, a curated event taxonomy, and conversion-like events with zero GA4 conversions/revenue.",{"type":47,"tag":48,"props":9297,"children":9298},{},[9299,9301,9307,9308,9314],{"type":52,"value":9300},"The caveat: GPT-5.5 hit repeated GA4 compatibility issues around traffic dimensions with ",{"type":47,"tag":332,"props":9302,"children":9304},{"className":9303},[],[9305],{"type":52,"value":9306},"eventCount",{"type":52,"value":7336},{"type":47,"tag":332,"props":9309,"children":9311},{"className":9310},[],[9312],{"type":52,"value":9313},"screenPageViews",{"type":52,"value":9315},", giving it the lowest accuracy score among successful Round 3 models.",{"type":47,"tag":1057,"props":9317,"children":9319},{"id":9318},"_9-glm-52",[9320],{"type":52,"value":9321},"9. GLM 5.2",{"type":47,"tag":48,"props":9323,"children":9324},{},[9325,9329,9331,9335,9337,9341,9342,9346],{"type":47,"tag":88,"props":9326,"children":9327},{},[9328],{"type":52,"value":3853},{"type":52,"value":9330}," $0.20 | ",{"type":47,"tag":88,"props":9332,"children":9333},{},[9334],{"type":52,"value":6872},{"type":52,"value":9336}," 84s | ",{"type":47,"tag":88,"props":9338,"children":9339},{},[9340],{"type":52,"value":6879},{"type":52,"value":9237},{"type":47,"tag":88,"props":9343,"children":9344},{},[9345],{"type":52,"value":6886},{"type":52,"value":7189},{"type":47,"tag":48,"props":9348,"children":9349},{},[9350],{"type":52,"value":9351},"GLM 5.2 correctly identified the dataset as synthetic with high confidence. It called out 100% engagement, zero new users, unassigned traffic, zero conversions, weekend traffic oddities, and generation-like documentation parameters.",{"type":47,"tag":48,"props":9353,"children":9354},{},[9355],{"type":52,"value":9356},"It was useful, but not as complete or as clean as the top models.",{"type":47,"tag":1057,"props":9358,"children":9360},{"id":9359},"_10-qwen37-plus",[9361],{"type":52,"value":9362},"10. Qwen3.7 Plus",{"type":47,"tag":48,"props":9364,"children":9365},{},[9366,9370,9371,9375,9377,9381,9382,9386],{"type":47,"tag":88,"props":9367,"children":9368},{},[9369],{"type":52,"value":3853},{"type":52,"value":4309},{"type":47,"tag":88,"props":9372,"children":9373},{},[9374],{"type":52,"value":6872},{"type":52,"value":9376}," 146s | ",{"type":47,"tag":88,"props":9378,"children":9379},{},[9380],{"type":52,"value":6879},{"type":52,"value":6881},{"type":47,"tag":88,"props":9383,"children":9384},{},[9385],{"type":52,"value":6886},{"type":52,"value":7277},{"type":47,"tag":48,"props":9388,"children":9389},{},[9390],{"type":52,"value":9391},"Qwen3.7 Plus had perfect field accuracy, but the final output was less evidence-rich than Qwen3.7 Max. It identified synthetic fingerprints, but did not provide the same level of audit depth.",{"type":47,"tag":48,"props":9393,"children":9394},{},[9395],{"type":52,"value":9396},"This is a good reminder that valid queries are not the same thing as a good analysis.",{"type":47,"tag":55,"props":9398,"children":9400},{"id":9399},"what-failed",[9401],{"type":52,"value":9402},"What Failed",{"type":47,"tag":48,"props":9404,"children":9405},{},[9406],{"type":52,"value":9407},"Claude Fable 5 was selected by the newest-model automation, but failed all 3 attempts through OpenRouter. Public reporting and Anthropic's own statement indicate Fable 5 had been disabled after a U.S. export-control directive, so we treated this as an access failure rather than a useful model-quality result and replaced it with GPT-5.5 as a one-off run.",{"type":47,"tag":48,"props":9409,"children":9410},{},[9411],{"type":52,"value":9412},"That replacement is now included in the Round 3 results.",{"type":47,"tag":55,"props":9414,"children":9416},{"id":9415},"demo-data-lessons",[9417],{"type":52,"value":9418},"Demo Data Lessons",{"type":47,"tag":48,"props":9420,"children":9421},{},[9422],{"type":52,"value":9423},"The models gave surprisingly useful feedback for improving synthetic analytics data.",{"type":47,"tag":48,"props":9425,"children":9426},{},[9427],{"type":52,"value":9428},"If the goal is a demo dataset that feels realistic without pretending to be production data, the biggest improvements are:",{"type":47,"tag":137,"props":9430,"children":9431},{},[9432,9437,9442,9447,9452,9457,9462],{"type":47,"tag":141,"props":9433,"children":9434},{},[9435],{"type":52,"value":9436},"Add long-tail geography instead of a small closed set of cities and countries.",{"type":47,"tag":141,"props":9438,"children":9439},{},[9440],{"type":52,"value":9441},"Add realistic browser and device tails: Edge, Firefox, Samsung Internet, bots, odd devices.",{"type":47,"tag":141,"props":9443,"children":9444},{},[9445],{"type":52,"value":9446},"Add seasonality: weekday/weekend dips, launch spikes, quiet periods, holidays.",{"type":47,"tag":141,"props":9448,"children":9449},{},[9450],{"type":52,"value":9451},"Add more acquisition mess: referrals, organic search, paid campaigns, direct traffic, spam.",{"type":47,"tag":141,"props":9453,"children":9454},{},[9455],{"type":52,"value":9456},"Add conversion inconsistencies that mirror real tracking: partial revenue, missing events, delayed conversions.",{"type":47,"tag":141,"props":9458,"children":9459},{},[9460],{"type":52,"value":9461},"Avoid perfect ratios and overly clean funnels.",{"type":47,"tag":141,"props":9463,"children":9464},{},[9465],{"type":52,"value":9466},"Keep explicit demo/test labeling so users are not misled.",{"type":47,"tag":48,"props":9468,"children":9469},{},[9470],{"type":52,"value":9471},"That last point matters. The goal is not to fool users. The goal is to create enough realism that the product, the AI, and the workflow are all tested honestly.",{"type":47,"tag":55,"props":9473,"children":9474},{"id":4603},[9475],{"type":52,"value":4606},{"type":47,"tag":48,"props":9477,"children":9478},{},[9479],{"type":47,"tag":88,"props":9480,"children":9481},{},[9482],{"type":52,"value":9483},"Test setup:",{"type":47,"tag":137,"props":9485,"children":9486},{},[9487,9492,9497,9502,9506,9511,9516,9521],{"type":47,"tag":141,"props":9488,"children":9489},{},[9490],{"type":52,"value":9491},"GA4 property ID: 509106858",{"type":47,"tag":141,"props":9493,"children":9494},{},[9495],{"type":52,"value":9496},"Same Anamap analytics benchmark runner as prior rounds",{"type":47,"tag":141,"props":9498,"children":9499},{},[9500],{"type":52,"value":9501},"Same system prompt and data-source tooling pattern",{"type":47,"tag":141,"props":9503,"children":9504},{},[9505],{"type":52,"value":6196},{"type":47,"tag":141,"props":9507,"children":9508},{},[9509],{"type":52,"value":9510},"Max 4 model turns per run",{"type":47,"tag":141,"props":9512,"children":9513},{},[9514],{"type":52,"value":9515},"OpenRouter model selection limited to models created in the last 3 months",{"type":47,"tag":141,"props":9517,"children":9518},{},[9519],{"type":52,"value":9520},"Previous benchmark models excluded",{"type":47,"tag":141,"props":9522,"children":9523},{},[9524],{"type":52,"value":9525},"GPT-5.5 added as a one-off replacement for Claude Fable 5 after access to Fable 5 was disabled following a U.S. export-control directive",{"type":47,"tag":48,"props":9527,"children":9528},{},[9529,9534],{"type":47,"tag":88,"props":9530,"children":9531},{},[9532],{"type":52,"value":9533},"Prompt:",{"type":52,"value":9535}," Determine whether the connected dataset is real, synthetic, or inconclusive, and explain the evidence.",{"type":47,"tag":48,"props":9537,"children":9538},{},[9539],{"type":47,"tag":88,"props":9540,"children":9541},{},[9542],{"type":52,"value":9543},"Evaluation criteria:",{"type":47,"tag":137,"props":9545,"children":9546},{},[9547,9552,9557,9562,9567,9571,9576,9581],{"type":47,"tag":141,"props":9548,"children":9549},{},[9550],{"type":52,"value":9551},"Completion rate across 3 runs",{"type":47,"tag":141,"props":9553,"children":9554},{},[9555],{"type":52,"value":9556},"Quality of final analysis",{"type":47,"tag":141,"props":9558,"children":9559},{},[9560],{"type":52,"value":9561},"GA4 field accuracy / hallucination score",{"type":47,"tag":141,"props":9563,"children":9564},{},[9565],{"type":52,"value":9566},"Whether the model queried connected data",{"type":47,"tag":141,"props":9568,"children":9569},{},[9570],{"type":52,"value":1485},{"type":47,"tag":141,"props":9572,"children":9573},{},[9574],{"type":52,"value":9575},"Latency",{"type":47,"tag":141,"props":9577,"children":9578},{},[9579],{"type":52,"value":9580},"Specificity of evidence",{"type":47,"tag":141,"props":9582,"children":9583},{},[9584],{"type":52,"value":9585},"Usefulness of demo-data improvement recommendations",{"type":47,"tag":48,"props":9587,"children":9588},{},[9589,9594],{"type":47,"tag":88,"props":9590,"children":9591},{},[9592],{"type":52,"value":9593},"Total cost:",{"type":52,"value":9595}," $4.62 across the initial 10-model run plus the GPT-5.5 replacement. The replacement leaderboard excludes the unavailable Claude Fable row.",{"type":47,"tag":55,"props":9597,"children":9599},{"id":9598},"what-this-means",[9600],{"type":52,"value":9601},"What This Means",{"type":47,"tag":48,"props":9603,"children":9604},{},[9605],{"type":52,"value":9606},"For synthetic-data detection, the best model is not necessarily the fastest or most expensive one.",{"type":47,"tag":48,"props":9608,"children":9609},{},[9610,9614,9616,9620,9622,9626,9628,9632],{"type":47,"tag":88,"props":9611,"children":9612},{},[9613],{"type":52,"value":1107},{"type":52,"value":9615}," was the best evidence-backed result. ",{"type":47,"tag":88,"props":9617,"children":9618},{},[9619],{"type":52,"value":5168},{"type":52,"value":9621}," was the fastest classifier. ",{"type":47,"tag":88,"props":9623,"children":9624},{},[9625],{"type":52,"value":1205},{"type":52,"value":9627}," gave the most useful low-cost nuanced answer. ",{"type":47,"tag":88,"props":9629,"children":9630},{},[9631],{"type":52,"value":5187},{"type":52,"value":9633}," was directionally right but expensive and less precise with GA4 query construction.",{"type":47,"tag":48,"props":9635,"children":9636},{},[9637],{"type":52,"value":9638},"The broader pattern is encouraging: most frontier models can identify synthetic analytics data when the evidence is available. The harder problem is making them show their work reliably.",{"type":47,"tag":1245,"props":9640,"children":9643},{"heading":9641,"icon":3338,"primary-link":1040,"primary-text":1251,"secondary-link":1249,"secondary-text":1250,"subheading":9642},"Compare all benchmark rounds","See every model from Round 1, Round 2, and Round 3 in the combined leaderboard.",[],{"type":47,"tag":55,"props":9645,"children":9646},{"id":1821},[9647],{"type":52,"value":1824},{"type":47,"tag":1057,"props":9649,"children":9651},{"id":9650},"which-model-was-best-at-detecting-synthetic-analytics-data",[9652],{"type":52,"value":9653},"Which model was best at detecting synthetic analytics data?",{"type":47,"tag":48,"props":9655,"children":9656},{},[9657],{"type":52,"value":9658},"Gemini 3.5 Flash was the best evidence-backed model in this benchmark. It completed all 3 runs, kept perfect GA4 field accuracy, made data requests, and produced a clear synthetic-data audit.",{"type":47,"tag":1057,"props":9660,"children":9662},{"id":9661},"why-did-grok-43-not-rank-first-if-it-was-fastest",[9663],{"type":52,"value":9664},"Why did Grok 4.3 not rank first if it was fastest?",{"type":47,"tag":48,"props":9666,"children":9667},{},[9668],{"type":52,"value":9669},"Grok 4.3 was the fastest successful model and produced a correct answer, but it made zero data requests. For a task that asks the model to use evidence from connected data sources, that matters.",{"type":47,"tag":1057,"props":9671,"children":9673},{"id":9672},"how-did-gpt-55-perform",[9674],{"type":52,"value":9675},"How did GPT-5.5 perform?",{"type":47,"tag":48,"props":9677,"children":9678},{},[9679],{"type":52,"value":9680},"GPT-5.5 completed all 3 runs and correctly classified the dataset as synthetic demo data. It averaged 148 seconds and $1.45 per run, with an 88/100 field-accuracy score due to GA4 compatibility issues.",{"type":47,"tag":1057,"props":9682,"children":9684},{"id":9683},"was-the-synthetic-data-too-obvious",[9685],{"type":52,"value":9686},"Was the synthetic data too obvious?",{"type":47,"tag":48,"props":9688,"children":9689},{},[9690],{"type":52,"value":9691},"Some signals were intentionally obvious, such as the test-property label. But the better models went beyond that and inspected distributions, attribution gaps, temporal coverage, event taxonomy, and conversion inconsistencies.",{"type":47,"tag":1057,"props":9693,"children":9695},{"id":9694},"should-demo-data-try-to-fool-ai-models",[9696],{"type":52,"value":9697},"Should demo data try to fool AI models?",{"type":47,"tag":48,"props":9699,"children":9700},{},[9701],{"type":52,"value":9702},"No. Demo data should be clearly labeled. The goal is not deception; it is realism. A good demo dataset should contain enough realistic messiness to test workflows and model judgment without misleading users.",{"type":47,"tag":1904,"props":9704,"children":9706},{":faqs":9705},"[{\"question\":\"Which model was best at detecting synthetic analytics data?\",\"answer\":\"Gemini 3.5 Flash was the best evidence-backed model in this benchmark. It completed all 3 runs, kept perfect GA4 field accuracy, made data requests, and produced a clear synthetic-data audit.\"},{\"question\":\"Why did Grok 4.3 not rank first if it was fastest?\",\"answer\":\"Grok 4.3 was the fastest successful model and produced a correct answer, but it made zero data requests. For a task that asks the model to use evidence from connected data sources, that matters.\"},{\"question\":\"How did GPT-5.5 perform?\",\"answer\":\"GPT-5.5 completed all 3 runs and correctly classified the dataset as synthetic demo data. It averaged 148 seconds and $1.45 per run, with an 88/100 field-accuracy score due to GA4 compatibility issues.\"},{\"question\":\"Was the synthetic data too obvious?\",\"answer\":\"Some signals were intentionally obvious, such as the test-property label. But the better models went beyond that and inspected distributions, attribution gaps, temporal coverage, event taxonomy, and conversion inconsistencies.\"},{\"question\":\"Should demo data try to fool AI models?\",\"answer\":\"No. Demo data should be clearly labeled. The goal is not deception; it is realism. A good demo dataset should contain enough realistic messiness to test workflows and model judgment without misleading users.\"}]",[9707,9711],{"type":47,"tag":55,"props":9708,"children":9709},{"id":1910},[9710],{"type":52,"value":1913},{"type":47,"tag":137,"props":9712,"children":9713},{},[9714,9723,9732,9741],{"type":47,"tag":141,"props":9715,"children":9716},{},[9717,9721],{"type":47,"tag":67,"props":9718,"children":9719},{"href":1040},[9720],{"type":52,"value":1924},{"type":52,"value":9722},": Combined leaderboard across all benchmark rounds",{"type":47,"tag":141,"props":9724,"children":9725},{},[9726,9730],{"type":47,"tag":67,"props":9727,"children":9728},{"href":912},[9729],{"type":52,"value":4926},{"type":52,"value":9731},": Recommendations by use case",{"type":47,"tag":141,"props":9733,"children":9734},{},[9735,9739],{"type":47,"tag":67,"props":9736,"children":9737},{"href":1329},[9738],{"type":52,"value":1944},{"type":52,"value":9740},": 6 newer models, 3 runs each",{"type":47,"tag":141,"props":9742,"children":9743},{},[9744,9748],{"type":47,"tag":67,"props":9745,"children":9746},{"href":1313},[9747],{"type":52,"value":1954},{"type":52,"value":9749},": The original attribution-quality benchmark",{"title":8,"searchDepth":587,"depth":588,"links":9751},[9752,9753,9754,9755,9756,9768,9769,9770,9771,9772,9779],{"id":8078,"depth":587,"text":8081},{"id":8194,"depth":587,"text":8197},{"id":6237,"depth":587,"text":6240},{"id":8860,"depth":587,"text":8863},{"id":6843,"depth":587,"text":6846,"children":9757},[9758,9759,9760,9761,9762,9763,9764,9765,9766,9767],{"id":8923,"depth":588,"text":8926},{"id":8996,"depth":588,"text":8999},{"id":9037,"depth":588,"text":9040},{"id":9082,"depth":588,"text":9085},{"id":9131,"depth":588,"text":9134},{"id":9173,"depth":588,"text":9176},{"id":9215,"depth":588,"text":9218},{"id":9255,"depth":588,"text":9258},{"id":9318,"depth":588,"text":9321},{"id":9359,"depth":588,"text":9362},{"id":9399,"depth":587,"text":9402},{"id":9415,"depth":587,"text":9418},{"id":4603,"depth":587,"text":4606},{"id":9598,"depth":587,"text":9601},{"id":1821,"depth":587,"text":1824,"children":9773},[9774,9775,9776,9777,9778],{"id":9650,"depth":588,"text":9653},{"id":9661,"depth":588,"text":9664},{"id":9672,"depth":588,"text":9675},{"id":9683,"depth":588,"text":9686},{"id":9694,"depth":588,"text":9697},{"id":1910,"depth":587,"text":1913},"content:blog:llm-analytics-benchmark-round-3-synthetic-data.md","blog/llm-analytics-benchmark-round-3-synthetic-data.md","blog/llm-analytics-benchmark-round-3-synthetic-data",{"_path":9784,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":9785,"description":9786,"draft":7,"publicationDate":9787,"updatedAt":9787,"image":9788,"author":9789,"head":9791,"category":9798,"tags":9799,"keywords":9794,"ogTitle":9785,"ogDescription":9804,"twitterTitle":9785,"twitterDescription":9805,"body":9806,"_type":596,"_id":9971,"_source":598,"_file":9972,"_stem":9973,"_extension":601},"/blog/product-data-interview-audris-wong","Product Data Interview: Audris Wong","Audris Wong, an independent consultant at Snowpack Data and former product manager at The Browser Company, shares how she approaches product decisions, measurement confidence, data access, and turning product data into action.","2026-06-17","/images/blog/product-manager-data-interviews-hero.webp",{"id":14,"name":15,"role":9790},"Founder & CEO",{"meta":9792},[9793,9795,9796,9797],{"name":33,"content":9794},"product manager interview, product analytics, data-driven product decisions, product metrics, Audris Wong, Snowpack Data, product management data",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},"Interviews",[9800,9801,9802,9803],"product management","product analytics","data-driven decisions","interviews","Audris Wong of Snowpack Data on product decisions, measurement confidence, self-serve data, and turning product analytics into action.","Audris Wong of Snowpack Data shares how she approaches product decisions with data.",{"type":44,"children":9807,"toc":9959},[9808,9813,9818,9829,9835,9840,9845,9851,9856,9862,9867,9872,9878,9883,9888,9894,9899,9905,9910,9916,9921,9927,9932,9938,9943,9949,9954],{"type":47,"tag":48,"props":9809,"children":9810},{},[9811],{"type":52,"value":9812},"This is part of our Product Data Interviews series, where we ask product managers across the industry the same set of questions about how they use data, what slows them down, and what helps them make better product decisions.",{"type":47,"tag":48,"props":9814,"children":9815},{},[9816],{"type":52,"value":9817},"Audris Wong is an independent consultant at Snowpack Data, where she helps clients evaluate production ML models, understand what is working, and simplify deployment without sacrificing signal. Before Snowpack Data, she worked at The Browser Company, where her product work focused heavily on onboarding: when to roll out new features, when to pull them back, and how to run hypothesis tests on new experiences designed to help more users get value from the product.",{"type":47,"tag":48,"props":9819,"children":9820},{},[9821,9828],{"type":47,"tag":67,"props":9822,"children":9825},{"href":9823,"rel":9824},"https://www.linkedin.com/in/audriswong",[71],[9826],{"type":52,"value":9827},"Connect with Audris on LinkedIn",{"type":52,"value":94},{"type":47,"tag":55,"props":9830,"children":9832},{"id":9831},"_1-what-kind-of-product-decisions-are-you-personally-responsible-for-in-your-day-to-day-work",[9833],{"type":52,"value":9834},"1. What kind of product decisions are you personally responsible for in your day to day work?",{"type":47,"tag":48,"props":9836,"children":9837},{},[9838],{"type":52,"value":9839},"At The Browser Company, a big part of my role was around the onboarding experience. Deciding when to roll out new features, when to pull them back, and running hypothesis tests on new experiences we were designing to improve how many users actually got value from the product and stuck around.",{"type":47,"tag":48,"props":9841,"children":9842},{},[9843],{"type":52,"value":9844},"Now in the consulting work I do, I help clients evaluate their production ML models - specifically figuring out how effective those models actually are, which parts were worth keeping, and how to simplify deployment without sacrificing signal.",{"type":47,"tag":55,"props":9846,"children":9848},{"id":9847},"_2-when-youre-evaluating-whether-something-is-working-or-worth-building-what-signals-matter-most-to-you",[9849],{"type":52,"value":9850},"2. When you're evaluating whether something is working (or worth building), what signals matter most to you?",{"type":47,"tag":48,"props":9852,"children":9853},{},[9854],{"type":52,"value":9855},"For figuring out if something is working, it really comes down to how it's performing against the metric we set out to move. For deciding if something's worth building, I try to be really intentional upfront. What is this feature actually designed to do? What does success look like? Once we have a clear answer to that, you can use it as a lens to evaluate how big the opportunity actually is and what the realistic upside is.",{"type":47,"tag":55,"props":9857,"children":9859},{"id":9858},"_3-walk-me-through-a-recent-product-decision-you-made-that-involved-data-what-did-the-process-actually-look-like",[9860],{"type":52,"value":9861},"3. Walk me through a recent product decision you made that involved data. What did the process actually look like?",{"type":47,"tag":48,"props":9863,"children":9864},{},[9865],{"type":52,"value":9866},"I start with answering the question \"what do we actually expect to learn from this?\". From there, I define what success and failure look like in concrete terms: if our hypothesis is right, what should we see in the data? If it's wrong, what would that look like? Then we make sure we're logging and tracking both the core metric we care about and any guardrail metrics that would tell us if something's going sideways.",{"type":47,"tag":48,"props":9868,"children":9869},{},[9870],{"type":52,"value":9871},"During data collection, I check in regularly and I'm really careful about not drawing conclusions too early. Once we have enough confidence in the data, we build out a decision matrix and make a recommendation that factors in the opportunity cost of each path. Then it goes to stakeholders - engineering, product, leadership - for a final review.",{"type":47,"tag":55,"props":9873,"children":9875},{"id":9874},"_4-how-confident-do-you-generally-feel-in-the-data-available-to-you-when-making-product-decisions-what-tends-to-increase-or-reduce-that-confidence",[9876],{"type":52,"value":9877},"4. How confident do you generally feel in the data available to you when making product decisions? What tends to increase or reduce that confidence?",{"type":47,"tag":48,"props":9879,"children":9880},{},[9881],{"type":52,"value":9882},"Honestly, I have pretty limited trust in existing data to measure exactly what I need since that's just the reality of most setups. What increases my confidence is being really explicit upfront: defining exactly what we need to measure, setting up new tracking specifically for that purpose, and not relying on something that was built for a different reason.",{"type":47,"tag":48,"props":9884,"children":9885},{},[9886],{"type":52,"value":9887},"I also try to explicitly outline where measurement might be off and what that would mean for the decision. Usually the risk of misalignment is smaller than it feels, but naming it out loud helps everyone calibrate.",{"type":47,"tag":55,"props":9889,"children":9891},{"id":9890},"_5-whats-the-most-frustrating-or-time-consuming-part-of-getting-the-insights-you-need-to-make-a-decision",[9892],{"type":52,"value":9893},"5. What's the most frustrating or time-consuming part of getting the insights you need to make a decision?",{"type":47,"tag":48,"props":9895,"children":9896},{},[9897],{"type":52,"value":9898},"Gaps in what's actually being tracked, and not being able to trust how things are defined. Without trust in logging, everything downstream in the decision process gets harder and slower.",{"type":47,"tag":55,"props":9900,"children":9902},{"id":9901},"_6-how-self-serve-is-data-access-for-product-managers-at-your-company-today",[9903],{"type":52,"value":9904},"6. How self-serve is data access for product managers at your company today?",{"type":47,"tag":48,"props":9906,"children":9907},{},[9908],{"type":52,"value":9909},"Fairly self-serviceable. We use Claude via MCP and Amplitude, so I can get to most of what I need without having to go through an analyst for every question.",{"type":47,"tag":55,"props":9911,"children":9913},{"id":9912},"_7-whats-the-hardest-thing-about-turning-data-into-action-rather-than-just-more-dashboards-or-reports",[9914],{"type":52,"value":9915},"7. What's the hardest thing about turning data into action rather than just more dashboards or reports?",{"type":47,"tag":48,"props":9917,"children":9918},{},[9919],{"type":52,"value":9920},"Getting enough context to move from \"what happened\" to \"why did it happen\" and \"what do we do about it.\" The data tells you something changed; figuring out the meaning behind it and what that means for your next move is where the real work is.",{"type":47,"tag":55,"props":9922,"children":9924},{"id":9923},"_8-are-there-product-metrics-or-definitions-that-people-at-the-company-regularly-interpret-differently",[9925],{"type":52,"value":9926},"8. Are there product metrics or definitions that people at the company regularly interpret differently?",{"type":47,"tag":48,"props":9928,"children":9929},{},[9930],{"type":52,"value":9931},"Conversion metrics are a classic one. There's a lot of room for creativity in how you define the numerator, the denominator, the time window. Two people can look at \"conversion\" and be talking about completely different things. That ambiguity causes a lot of friction.",{"type":47,"tag":55,"props":9933,"children":9935},{"id":9934},"_9-whats-a-surprising-or-overlooked-source-of-product-insight-that-you-think-more-teams-should-pay-attention-to",[9936],{"type":52,"value":9937},"9. What's a surprising or overlooked source of product insight that you think more teams should pay attention to?",{"type":47,"tag":48,"props":9939,"children":9940},{},[9941],{"type":52,"value":9942},"The opportunity cost of not making a decision. Teams spend so much time trying to get more certainty before acting, but if you frame the cost of inaction clearly, it really helps cut through analysis paralysis. Sometimes waiting is the most expensive option on the table.",{"type":47,"tag":55,"props":9944,"children":9946},{"id":9945},"_10-what-advice-would-you-give-another-pm-at-a-startup-trying-to-make-better-product-decisions-with-data",[9947],{"type":52,"value":9948},"10. What advice would you give another PM at a startup trying to make better product decisions with data?",{"type":47,"tag":48,"props":9950,"children":9951},{},[9952],{"type":52,"value":9953},"Invest time in the basics. Get clear, shared definitions for the concepts that matter across your business, things like what counts as an active user, what activation means, what engagement looks like. It's not glamorous work, but the return on it is enormous.",{"type":47,"tag":48,"props":9955,"children":9956},{},[9957],{"type":52,"value":9958},"When these things aren't nailed down, they can become blockers when you need to move fast and make a decision. Getting ahead of that early saves a lot of pain later.",{"title":8,"searchDepth":587,"depth":588,"links":9960},[9961,9962,9963,9964,9965,9966,9967,9968,9969,9970],{"id":9831,"depth":587,"text":9834},{"id":9847,"depth":587,"text":9850},{"id":9858,"depth":587,"text":9861},{"id":9874,"depth":587,"text":9877},{"id":9890,"depth":587,"text":9893},{"id":9901,"depth":587,"text":9904},{"id":9912,"depth":587,"text":9915},{"id":9923,"depth":587,"text":9926},{"id":9934,"depth":587,"text":9937},{"id":9945,"depth":587,"text":9948},"content:blog:product-data-interview-audris-wong.md","blog/product-data-interview-audris-wong.md","blog/product-data-interview-audris-wong",{"_path":9975,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":9976,"description":9977,"draft":7,"publicationDate":9978,"updatedAt":9978,"image":9788,"author":9979,"head":9980,"category":9798,"tags":9987,"keywords":9983,"ogTitle":9976,"ogDescription":9988,"twitterTitle":9976,"twitterDescription":9989,"body":9990,"_type":596,"_id":10223,"_source":598,"_file":10224,"_stem":10225,"_extension":601},"/blog/product-data-interview-making-better-decisions-with-data","Product Data Interview: Making Better Decisions With Data","An experienced product manager shares how segmentation, qualitative research, trustworthy instrumentation, and healthy skepticism lead to better product decisions.","2026-07-20",{"id":14,"name":15,"role":9790},{"meta":9981},[9982,9984,9985,9986],{"name":33,"content":9983},"product manager interview, product analytics, data-driven product decisions, product metrics, data quality, experimentation, customer segmentation",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},[9800,9801,9802,9803],"An experienced PM on segmentation, trustworthy instrumentation, experimentation, and challenging the story behind the data.","An experienced product manager shares practical lessons for making better product decisions with data.",{"type":44,"children":9991,"toc":10210},[9992,9996,10001,10006,10012,10017,10022,10027,10032,10037,10041,10046,10051,10055,10060,10065,10070,10075,10079,10084,10089,10093,10098,10103,10108,10113,10117,10122,10127,10132,10137,10141,10146,10151,10156,10160,10165,10170,10174,10179,10184,10189,10195,10200,10205],{"type":47,"tag":48,"props":9993,"children":9994},{},[9995],{"type":52,"value":9812},{"type":47,"tag":48,"props":9997,"children":9998},{},[9999],{"type":52,"value":10000},"Our interviewee is an experienced product manager with a prior background in software development at a consumer technology company. Their work spanned product discovery, measurement, experimentation, user-interface improvements, and monetization.",{"type":47,"tag":48,"props":10002,"children":10003},{},[10004],{"type":52,"value":10005},"This interview has been lightly edited for clarity and anonymized. Company and product names, internal organizations, career timelines, and product-specific operational details have been removed or generalized.",{"type":47,"tag":55,"props":10007,"children":10009},{"id":10008},"_1-what-kind-of-product-decisions-were-you-personally-responsible-for-in-your-previous-role",[10010],{"type":52,"value":10011},"1. What kind of product decisions were you personally responsible for in your previous role?",{"type":47,"tag":48,"props":10013,"children":10014},{},[10015],{"type":52,"value":10016},"I worked across consumer-facing experiences, data instrumentation, user-interface redesigns, monetization, testing, and experimentation. I also worked on systems that helped people discover relevant options and supported broader product-integration work.",{"type":47,"tag":48,"props":10018,"children":10019},{},[10020],{"type":52,"value":10021},"My background helped me work across both the technical and product sides because I worked as a developer before moving into product management.",{"type":47,"tag":55,"props":10023,"children":10024},{"id":9847},[10025],{"type":52,"value":10026},"2. When you're evaluating whether something is working or worth building, what signals matter most to you?",{"type":47,"tag":48,"props":10028,"children":10029},{},[10030],{"type":52,"value":10031},"I prefer a blend of quantitative behavioral data and qualitative user research. Behavioral data matters, but so do direct customer quotes, bug reports, and even emotionally worded feedback. When a customer cares enough to send a detailed complaint, there is often a useful signal underneath the intensity. The exact severity may not generalize to every customer, but the underlying problem can.",{"type":47,"tag":48,"props":10033,"children":10034},{},[10035],{"type":52,"value":10036},"I also think companies can over-focus on whatever leadership has labeled the highest priority. Teams need room to build small \"delighter\" features as well. A product people enjoy can earn more patience when other parts of the experience are imperfect. In an ideal setup, teams would reserve some capacity for promising small ideas.",{"type":47,"tag":55,"props":10038,"children":10039},{"id":9858},[10040],{"type":52,"value":9861},{"type":47,"tag":48,"props":10042,"children":10043},{},[10044],{"type":52,"value":10045},"We were evaluating a product change intended to serve a particular customer segment better. Rather than use company-wide averages, we asked data scientists to isolate that cohort and analyze its specific behavior. The goal was to understand what that segment valued and design an option that would actually fit its needs.",{"type":47,"tag":48,"props":10047,"children":10048},{},[10049],{"type":52,"value":10050},"The main lesson was that averages can hide a collection of dissimilar customers whose behaviors may even conflict. Good product analysis starts by identifying meaningful customer groups, learning what each expects, and then evaluating an option against the relevant cohort instead of an undifferentiated average. User research is essential because it helps define those groupings before the quantitative analysis begins.",{"type":47,"tag":55,"props":10052,"children":10053},{"id":9874},[10054],{"type":52,"value":9877},{"type":47,"tag":48,"props":10056,"children":10057},{},[10058],{"type":52,"value":10059},"Access to dedicated data science and user-research teams increased my confidence. Data scientists could design more sophisticated analyses and challenge assumptions in the numbers, while researchers could provide the qualitative context needed to understand why customers behaved a certain way. Having both perspectives made it easier to cross-check a conclusion instead of relying on a single signal.",{"type":47,"tag":48,"props":10061,"children":10062},{},[10063],{"type":52,"value":10064},"In practice, though, my confidence was often low because those teams were spread too thin and frequently redirected by ad hoc executive requests. A PM might wait weeks or months for analysis, and user research could take even longer to reach the front of the backlog. That made it difficult to get evidence on the timeline of a real product decision.",{"type":47,"tag":48,"props":10066,"children":10067},{},[10068],{"type":52,"value":10069},"The dashboards were not always maintained reliably, either. For example, a filter could appear to work while still returning unfiltered results. Unless someone noticed that the numbers looked suspicious and asked the owner, a team could make a decision from the wrong slice of data.",{"type":47,"tag":48,"props":10071,"children":10072},{},[10073],{"type":52,"value":10074},"Confidence also increases when the underlying pipeline is maintained, filters are validated, definitions are visible, the result can be cross-checked against another source, and the data scientist or researcher has enough time to understand the question. It decreases when data is passed through several layers of storytelling, when the source or agenda is unclear, or when teams are pressured to find a favorable interpretation.",{"type":47,"tag":55,"props":10076,"children":10077},{"id":9890},[10078],{"type":52,"value":9893},{"type":47,"tag":48,"props":10080,"children":10081},{},[10082],{"type":52,"value":10083},"The hardest part was combining data across systems. Some information had to be separated by design to prevent teams from seeing competitively sensitive data. Restricting that data was the correct and ethical policy. At the same time, demonstrating a valid business reason and securing approval made access personally time-consuming.",{"type":47,"tag":48,"props":10085,"children":10086},{},[10087],{"type":52,"value":10088},"Compounding that, different parts of the customer journey were measured through several analytics systems rather than one end-to-end pipeline. Those systems used incompatible collection methods, which meant their numbers could not simply be joined. A PM might understand one part of the journey but be unable to connect it reliably to another. On top of that, we still had to wait for the data science team to prioritize the request.",{"type":47,"tag":55,"props":10090,"children":10091},{"id":9901},[10092],{"type":52,"value":9904},{"type":47,"tag":48,"props":10094,"children":10095},{},[10096],{"type":52,"value":10097},"Basic dashboard exploration was self-serve once a PM had been granted permission. We could use prebuilt dashboards to pull common insights without asking an analyst to answer every question.",{"type":47,"tag":48,"props":10099,"children":10100},{},[10101],{"type":52,"value":10102},"The limits appeared when we needed to create a new analysis, validate a dashboard that might be broken, join data from different systems, or do more sophisticated work. That was not self-serve. It required dedicated data scientists, and they generally produced deeper and more reliable insights than a PM could get from the dashboards alone.",{"type":47,"tag":48,"props":10104,"children":10105},{},[10106],{"type":52,"value":10107},"The company also experimented with using LLM-based generative AI to produce dashboards and write-ups from datasets. Those tools sometimes hallucinated figures that did not reflect the underlying data.",{"type":47,"tag":48,"props":10109,"children":10110},{},[10111],{"type":52,"value":10112},"The problem was not automation, or even AI more broadly. It was applying LLMs to tasks they were poorly suited for without enough validation. The team had not drawn a clear enough boundary between work suited to an LLM and work better handled through deterministic automation, traditional machine learning, or dedicated data scientists. Giving more people access is useful only if the system also provides trustworthy definitions, lineage, validation, and guardrails.",{"type":47,"tag":55,"props":10114,"children":10115},{"id":9912},[10116],{"type":52,"value":9915},{"type":47,"tag":48,"props":10118,"children":10119},{},[10120],{"type":52,"value":10121},"The hardest part was deciding what outcome actually represented customer success, especially when a company-level metric pointed in the opposite direction. Time spent finding something is a good example. A company might treat more time in a product as better engagement, but the purpose of a discovery experience is to help customers find what they need quickly. If people spend less time searching because the product helped them complete their intended action sooner, the lower number may represent a better customer outcome.",{"type":47,"tag":48,"props":10123,"children":10124},{},[10125],{"type":52,"value":10126},"There was also pressure to turn every result into a positive story. In one feature experiment, the interaction volume was suspiciously low while completion among those few users appeared perfect. Leadership wanted to emphasize the completion rate. The more likely explanation was that internal testers were overrepresented and instrumentation for real customers was not working correctly.",{"type":47,"tag":48,"props":10128,"children":10129},{},[10130],{"type":52,"value":10131},"We also found unexpected imbalances between treatment and control populations after certain filters were applied, which could invalidate the result. But identifying that problem could be treated as obstruction, while a favorable slice was easier to present.",{"type":47,"tag":48,"props":10133,"children":10134},{},[10135],{"type":52,"value":10136},"Turning data into action requires an environment where teams can say, \"This experiment is inconclusive,\" discard invalid results, and investigate instrumentation before making a decision. Storytelling should make sound evidence understandable; it should not convert weak evidence into certainty.",{"type":47,"tag":55,"props":10138,"children":10139},{"id":9923},[10140],{"type":52,"value":9926},{"type":47,"tag":48,"props":10142,"children":10143},{},[10144],{"type":52,"value":10145},"Yes. Engagement time was a common example. If average engagement falls, that could reflect a product change, a discovery problem introduced by a redesign, an outage, seasonality, or a change in customer routines. The metric describes the shape of the outcome, but it does not identify the cause.",{"type":47,"tag":48,"props":10147,"children":10148},{},[10149],{"type":52,"value":10150},"People also treated time spent as a clean measure of preference. That overlooks differences in experience length, shared use, activity elsewhere, accidental launches, one-time trials, and changing tastes. A customer can spend a long time with something without valuing it most, or complete a short experience and value it highly. Explicit feedback and implicit behavioral data answer different questions and should be used together.",{"type":47,"tag":48,"props":10152,"children":10153},{},[10154],{"type":52,"value":10155},"Experiment results were also interpreted differently depending on sample size and duration. Small expected effects require larger samples or longer tests, and some product effects need time to stabilize. Shipping pressure can make teams want an answer before the experiment has enough statistical power to provide one.",{"type":47,"tag":55,"props":10157,"children":10158},{"id":9934},[10159],{"type":52,"value":9937},{"type":47,"tag":48,"props":10161,"children":10162},{},[10163],{"type":52,"value":10164},"Public conversations on social media, forums, and community sites are often dismissed as a vocal minority, but they can be extremely useful. I once grouped a broad set of public comments into recurring themes. That made the feedback much more actionable than a handful of isolated screenshots.",{"type":47,"tag":48,"props":10166,"children":10167},{},[10168],{"type":52,"value":10169},"The important thing is not to interpret every comment literally. Look for the emotion and motivation underneath it. If many people express frustration or distrust, ask which unmet expectations could have produced that feeling. Public discussion is organic evidence of how customers frame the product in their own language, and one person who comments may represent many others who feel the same way but stay silent.",{"type":47,"tag":55,"props":10171,"children":10172},{"id":9945},[10173],{"type":52,"value":9948},{"type":47,"tag":48,"props":10175,"children":10176},{},[10177],{"type":52,"value":10178},"Always ask whether you are seeing correlation or causation, and immediately brainstorm the confounding variables that could create the same data shape. Make it a habit—or even a team exercise—to generate as many plausible explanations as possible. An explanation that sounds unlikely at first can turn out to be the real one.",{"type":47,"tag":48,"props":10180,"children":10181},{},[10182],{"type":52,"value":10183},"Question the data itself, too. Bad data is worse than no data because it creates unearned confidence. For example, an average-engagement metric can look absurdly low if it includes accidental launches and people who sampled an experience briefly and left. If that noise is not filtered, a team could draw the wrong conclusion about the product's value.",{"type":47,"tag":48,"props":10185,"children":10186},{},[10187],{"type":52,"value":10188},"Cross-reference different sources, validate instrumentation, inspect filters and denominators, and make sure the measurement matches the product behavior you are trying to understand. With no data, people tend to be cautious. With bad data, they can make a seriously damaging product decision with complete confidence.",{"type":47,"tag":55,"props":10190,"children":10192},{"id":10191},"editors-note-why-these-lessons-matter-to-anamap",[10193],{"type":52,"value":10194},"Editor's note: Why these lessons matter to Anamap",{"type":47,"tag":48,"props":10196,"children":10197},{},[10198],{"type":52,"value":10199},"Several themes in this conversation reflect why we created Anamap. Product teams should not have to make important decisions without knowing whether their instrumentation is reliable, what a metric means, or where the underlying data came from. Greater access to data only helps when it is paired with shared definitions, context, and ways to validate what the numbers appear to show.",{"type":47,"tag":48,"props":10201,"children":10202},{},[10203],{"type":52,"value":10204},"Anamap is our attempt to make that foundation more visible and easier to maintain. The goal is not to remove healthy skepticism from product decisions, but to give teams enough context to ask better questions, catch measurement problems earlier, and use data with more confidence.",{"type":47,"tag":48,"props":10206,"children":10207},{},[10208],{"type":52,"value":10209},"This note reflects our perspective on the lessons from the conversation and should not be read as an endorsement of Anamap by the interviewee.",{"title":8,"searchDepth":587,"depth":588,"links":10211},[10212,10213,10214,10215,10216,10217,10218,10219,10220,10221,10222],{"id":10008,"depth":587,"text":10011},{"id":9847,"depth":587,"text":10026},{"id":9858,"depth":587,"text":9861},{"id":9874,"depth":587,"text":9877},{"id":9890,"depth":587,"text":9893},{"id":9901,"depth":587,"text":9904},{"id":9912,"depth":587,"text":9915},{"id":9923,"depth":587,"text":9926},{"id":9934,"depth":587,"text":9937},{"id":9945,"depth":587,"text":9948},{"id":10191,"depth":587,"text":10194},"content:blog:product-data-interview-making-better-decisions-with-data.md","blog/product-data-interview-making-better-decisions-with-data.md","blog/product-data-interview-making-better-decisions-with-data",{"_path":10227,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":10228,"description":10229,"draft":7,"publicationDate":10230,"updatedAt":10230,"image":9788,"author":10231,"head":10232,"category":9798,"tags":10239,"keywords":10235,"ogTitle":10228,"ogDescription":10240,"twitterTitle":10228,"twitterDescription":10241,"body":10242,"_type":596,"_id":10481,"_source":598,"_file":10482,"_stem":10483,"_extension":601},"/blog/product-data-interview-randy-young","Product Data Interview: Randy Young","Randy Young, a forward-deployed product manager working on voice AI agents, shares how he approaches product decisions, measurement ambiguity, data access, and customer outcomes.","2026-06-23",{"id":14,"name":15,"role":9790},{"meta":10233},[10234,10236,10237,10238],{"name":33,"content":10235},"product manager interview, product analytics, data-driven product decisions, product metrics, Randy Young, voice AI, product management data",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":41},[9800,9801,9802,9803],"Randy Young on product decisions, voice AI measurement, data access, and turning customer call data into action.","Randy Young shares how he approaches product decisions with data.",{"type":44,"children":10243,"toc":10468},[10244,10248,10253,10264,10272,10276,10281,10286,10291,10296,10300,10305,10310,10314,10319,10324,10329,10333,10338,10343,10347,10352,10357,10361,10366,10371,10376,10380,10385,10389,10394,10399,10404,10409,10414,10418,10423,10428,10432,10437,10442,10447,10452,10457,10463],{"type":47,"tag":48,"props":10245,"children":10246},{},[10247],{"type":52,"value":9812},{"type":47,"tag":48,"props":10249,"children":10250},{},[10251],{"type":52,"value":10252},"Randy Young is a forward-deployed product manager at Cresta, where he works on voice AI agents that replace interactive voice response systems. He has previously worked in product roles at Autodesk, Splunk, Amazon, and Bugcrowd, with experience spanning customer-facing product management, product-led growth, experimentation, and enterprise customer workflows.",{"type":47,"tag":48,"props":10254,"children":10255},{},[10256,10263],{"type":47,"tag":67,"props":10257,"children":10260},{"href":10258,"rel":10259},"https://www.linkedin.com/in/randallyoung/",[71],[10261],{"type":52,"value":10262},"Connect with Randy on LinkedIn",{"type":52,"value":94},{"type":47,"tag":10265,"props":10266,"children":10271},"prose-img",{":height":10267,":width":10268,"alt":10269,"src":10270},"300","240","Randy Young","https://anamap.blob.core.windows.net/images/blog/headshots/randall-young-headshot.webp",[],{"type":47,"tag":55,"props":10273,"children":10274},{"id":9831},[10275],{"type":52,"value":9834},{"type":47,"tag":48,"props":10277,"children":10278},{},[10279],{"type":52,"value":10280},"My role at Cresta is very different from anything I have done before. Previously, I worked as a core, customer-facing product manager with a focus on product-led growth. That work was very much about experimentation: building features, testing outcomes, and using the results to guide the roadmap.",{"type":47,"tag":48,"props":10282,"children":10283},{},[10284],{"type":52,"value":10285},"Today I am a forward-deployed product manager building voice AI agents that replace IVRs. It is a little more like being a consultant product manager for customers. In many cases, you are working with call center operations teams that do not have a product manager, so you are helping them refine requirements, manage the project, and run user acceptance testing.",{"type":47,"tag":48,"props":10287,"children":10288},{},[10289],{"type":52,"value":10290},"Most of the decisions are around setting up a voice agent so it can meet the business goals the customer has. That could mean deflecting minutes from human agents through authentication, intent routing, or answering simple questions with an AI agent, such as account balances. It could also mean warming leads for sales or containing users completely when a human decision or human help is not necessary.",{"type":47,"tag":48,"props":10292,"children":10293},{},[10294],{"type":52,"value":10295},"The interesting thing about this position is that the product changes based on each customer's needs, which means the decisions are changing constantly as well.",{"type":47,"tag":55,"props":10297,"children":10298},{"id":9847},[10299],{"type":52,"value":9850},{"type":47,"tag":48,"props":10301,"children":10302},{},[10303],{"type":52,"value":10304},"You can measure a lot of things inside a call, especially the successful outcomes of the use cases you are trying to solve.",{"type":47,"tag":48,"props":10306,"children":10307},{},[10308],{"type":52,"value":10309},"There are also downstream effects you can measure, like a reduction in call abandonment, containment of the user inside the AI agent after the call is successfully handled, and a reduction in mean time to handle for the human agent after the call is handed off from the AI agent.",{"type":47,"tag":55,"props":10311,"children":10312},{"id":9858},[10313],{"type":52,"value":9861},{"type":47,"tag":48,"props":10315,"children":10316},{},[10317],{"type":52,"value":10318},"I had a customer that did not want to contain calls inside the voice AI agent because they thought their human agents would handle things better. At the same time, they had a problem with long hold times because their human agents were handling too many calls, and they had a high abandonment rate on calls that were put on hold.",{"type":47,"tag":48,"props":10320,"children":10321},{},[10322],{"type":52,"value":10323},"We loaded a few thousand of their human-agent calls into our system. Using AI, I evaluated those calls to see which questions could be answered easily and which issues could be solved inside a voice agent without ever needing to go to a human. I also evaluated the mean time to handle for each call.",{"type":47,"tag":48,"props":10325,"children":10326},{},[10327],{"type":52,"value":10328},"With that analysis, we could recommend a set of FAQ questions and knowledge-based articles that would solve some of their calls without requiring a human agent, while freeing up their human agents to solve the more complicated customer issues.",{"type":47,"tag":55,"props":10330,"children":10331},{"id":9874},[10332],{"type":52,"value":9877},{"type":47,"tag":48,"props":10334,"children":10335},{},[10336],{"type":52,"value":10337},"Sometimes you have a challenge getting to historical calls, and you have to make a judgment call based on other projects you have worked on.",{"type":47,"tag":48,"props":10339,"children":10340},{},[10341],{"type":52,"value":10342},"In those cases, you really cannot measure it until you launch, experiment with changes, and then measure the outcome.",{"type":47,"tag":55,"props":10344,"children":10345},{"id":9890},[10346],{"type":52,"value":9893},{"type":47,"tag":48,"props":10348,"children":10349},{},[10350],{"type":52,"value":10351},"In many cases, it can be challenging to get customers on board with making changes to their existing processes. They have a lot of experience running their businesses, but not a whole lot of experience with voice AI agents.",{"type":47,"tag":48,"props":10353,"children":10354},{},[10355],{"type":52,"value":10356},"In some cases, customers are not comfortable making decisions in ambiguity or with a limited amount of data.",{"type":47,"tag":55,"props":10358,"children":10359},{"id":9901},[10360],{"type":52,"value":9904},{"type":47,"tag":48,"props":10362,"children":10363},{},[10364],{"type":52,"value":10365},"Out of all the places I have worked, we probably have the most access to information about the performance of our product.",{"type":47,"tag":48,"props":10367,"children":10368},{},[10369],{"type":52,"value":10370},"I think it is important to have a combination of aggregated data across multiple customers, the ability to filter by industry or specific customer type or size, and a process for measuring specific customers and outcomes based on their AI agent's performance.",{"type":47,"tag":48,"props":10372,"children":10373},{},[10374],{"type":52,"value":10375},"It is also nice to have AI overlays that can give you additional insights.",{"type":47,"tag":55,"props":10377,"children":10378},{"id":9912},[10379],{"type":52,"value":9915},{"type":47,"tag":48,"props":10381,"children":10382},{},[10383],{"type":52,"value":10384},"I think this comes with experience and domain knowledge. You get more comfortable working in some ambiguity because you have some idea about how different decisions will affect things downstream.",{"type":47,"tag":55,"props":10386,"children":10387},{"id":9923},[10388],{"type":52,"value":9926},{"type":47,"tag":48,"props":10390,"children":10391},{},[10392],{"type":52,"value":10393},"Yes. There are some challenges around what is considered containment inside a voice agent.",{"type":47,"tag":48,"props":10395,"children":10396},{},[10397],{"type":52,"value":10398},"For example, if someone does not complete a task but gets an FAQ question answered and then hangs up, that customer did not get to a human agent, so they were contained. You have to make a judgment call on whether that is acceptable containment. Some customers only consider it containment when a happy path was completed, like making a payment or getting account information.",{"type":47,"tag":48,"props":10400,"children":10401},{},[10402],{"type":52,"value":10403},"There are also judgment calls around cases where someone gets a balance, their next payment due date, and the amount of that payment, but then hangs up before saying \"thank you.\" In many of these cases, you need to decide whether that counts as containment.",{"type":47,"tag":48,"props":10405,"children":10406},{},[10407],{"type":52,"value":10408},"Every product manager in our organization handles these differently, and it really depends on what the customer wants or thinks are acceptable outcomes.",{"type":47,"tag":48,"props":10410,"children":10411},{},[10412],{"type":52,"value":10413},"Recently we have been looking into first call resolution, or seeing if a customer calls back in three to five days for the same or a similar issue. Measuring this is more about true containment and a completely solved inquiry.",{"type":47,"tag":55,"props":10415,"children":10416},{"id":9934},[10417],{"type":52,"value":9937},{"type":47,"tag":48,"props":10419,"children":10420},{},[10421],{"type":52,"value":10422},"In every company I have worked at, it has always been a challenge to aggregate all of the different venues where a customer might touch your brand.",{"type":47,"tag":48,"props":10424,"children":10425},{},[10426],{"type":52,"value":10427},"It is about the single view of your customer: understanding all the different avenues where they have touched your company and the success of those different avenues. That could include your support organization, social media, a conference, sales outreach, a chat or phone system, talking to a human support agent, or even their satisfaction with your end product in general.",{"type":47,"tag":55,"props":10429,"children":10430},{"id":9945},[10431],{"type":52,"value":9948},{"type":47,"tag":48,"props":10433,"children":10434},{},[10435],{"type":52,"value":10436},"Initially, especially when you are building a zero-to-one product, you need to listen to your customer. You need to do lots of interviews, listen for signals, and then build your product based on those signals.",{"type":47,"tag":48,"props":10438,"children":10439},{},[10440],{"type":52,"value":10441},"As you are building, you need lighthouse customers to bounce ideas off and measure the success of what you are building and prototyping. It is really important to define success before you even start building and to have a way to measure against that.",{"type":47,"tag":48,"props":10443,"children":10444},{},[10445],{"type":52,"value":10446},"Success can be revenue. Success can be customer adoption. Success can even be customer satisfaction based on use of the product. It really depends on the product you are building and the intended outcome.",{"type":47,"tag":48,"props":10448,"children":10449},{},[10450],{"type":52,"value":10451},"For example, enterprise software is usually about revenue or the number of Fortune 500 companies that are referenceable, which I like to call your \"NASCAR slide.\" Consumer products are usually about active users, subscribers, or ad revenue.",{"type":47,"tag":48,"props":10453,"children":10454},{},[10455],{"type":52,"value":10456},"It is all about setting realistic goals and then measuring, refining, and remeasuring your progress.",{"type":47,"tag":55,"props":10458,"children":10460},{"id":10459},"a-final-irony-about-data-companies",[10461],{"type":52,"value":10462},"A final irony about data companies",{"type":47,"tag":48,"props":10464,"children":10465},{},[10466],{"type":52,"value":10467},"One irony Randy called out is that companies considered \"data\" companies are sometimes the places where internal data is hardest to get to. Or, as he put it: \"The cobbler's kids have the worst shoes.\"",{"title":8,"searchDepth":587,"depth":588,"links":10469},[10470,10471,10472,10473,10474,10475,10476,10477,10478,10479,10480],{"id":9831,"depth":587,"text":9834},{"id":9847,"depth":587,"text":9850},{"id":9858,"depth":587,"text":9861},{"id":9874,"depth":587,"text":9877},{"id":9890,"depth":587,"text":9893},{"id":9901,"depth":587,"text":9904},{"id":9912,"depth":587,"text":9915},{"id":9923,"depth":587,"text":9926},{"id":9934,"depth":587,"text":9937},{"id":9945,"depth":587,"text":9948},{"id":10459,"depth":587,"text":10462},"content:blog:product-data-interview-randy-young.md","blog/product-data-interview-randy-young.md","blog/product-data-interview-randy-young",{"_path":10485,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":10486,"description":10487,"draft":7,"publicationDate":10488,"image":10489,"author":10490,"head":10491,"category":733,"body":10498,"_type":596,"_id":10575,"_source":598,"_file":10576,"_stem":10577,"_extension":601},"/blog/update-legacy-kpis-and-processes","Always Scrutinize Legacy KPIs and Processes","Legacy KPIs and processes can seem benign enough by only reducing efficiency by 2% but enough of them can compound together to create huge inefficiency.","2024-06-27","/images/blog/902d8e-87f4-446c-9dd3-8672c057db4d.webp",{"id":14,"name":15,"role":16},{"meta":10492},[10493,10495,10496,10497],{"name":33,"content":10494},"KPIs, data, business processes, business strategy",{"name":35,"content":36},{"name":38,"content":15},{"name":40,"content":616},{"type":44,"children":10499,"toc":10569},[10500,10506,10511,10516,10522,10534,10539,10545,10550,10556],{"type":47,"tag":55,"props":10501,"children":10503},{"id":10502},"what-are-legacy-kpis",[10504],{"type":52,"value":10505},"What are Legacy KPIs?",{"type":47,"tag":48,"props":10507,"children":10508},{},[10509],{"type":52,"value":10510},"Let’s start by explaining key performance indicators (KPIs) so we can set a baseline. KPIs are what businesses use to measure their performance, typically across time but they can also be used for one-time analysis. There’s a semi-common turn of phrase in the professional world that states “if you can’t measure it, you can’t improve it”. If you don’t have visibility into how your actions can positively or negatively impact a certain KPI there is no way to systematically improve it. As such, these key performance indicators serve as a centerpiece for organizations to make decisions about how to change course or how to make revenue generating improvements (ideally).",{"type":47,"tag":48,"props":10512,"children":10513},{},[10514],{"type":52,"value":10515},"Legacy KPIs specifically are KPIs that are old. Now what do I mean by old? There’s a little wiggle room in the interpretation because it depends on the speed of your businesses transitions and the speed of market / customer changes. For an average internet-based business anything that was initially conceived more than 2 years ago could be considered legacy. You might be saying “what!? 2 years is so short”. It is, I agree with you; on the flip side the world that many businesses exist in can change dramatically in two years. Think about any of the annual review cycles you’ve experienced. How often are the goals that are set at the beginning of the year the same as the true focus of the company at the end of the year? Rarely, if ever, right? If the priorities for your own role shift enough within a year it should make sense that company level measurement might also get stale after 2 years.",{"type":47,"tag":55,"props":10517,"children":10519},{"id":10518},"how-to-adapt-kpis-in-a-changing-world",[10520],{"type":52,"value":10521},"How to Adapt KPIs in a Changing World",{"type":47,"tag":48,"props":10523,"children":10524},{},[10525,10527,10532],{"type":52,"value":10526},"There are two main levels of adapting and re-assessing legacy measurement practices. The first and most obvious version for most people on a day-to-day basis is whether your KPIs capture your current business objectives. By that I mean if your company is leaning more into engagement on a specific section or vertical of your business your metric prioritization should reflect that. If you’re a legacy e-commerce platform and your growth has plateaued it might make sense to start incorporating more KPIs about your email list churn rate or your re-marketing channels. In summary, make sure you’re tracking KPIs that align with your ",{"type":47,"tag":88,"props":10528,"children":10529},{},[10530],{"type":52,"value":10531},"current",{"type":52,"value":10533}," business goals.",{"type":47,"tag":48,"props":10535,"children":10536},{},[10537],{"type":52,"value":10538},"The second level of adapting is re-evaluating the way that your existing metrics are calculated if those metrics meet the criteria of “this is relevant to my current business goals”. Companies rarely ever implement metrics perfectly the first time. Frequently, things are left half implemented or implemented in a workaround way where what is being measured isn’t the thing that business directly cares about. As an example, imagine your organization is a publisher for advertisements and is therefore interested in the number of ad impressions on specific types of pages. Maybe tracking the number of impressions on a given page wasn’t possible when the need first arose so instead of tracking impressions your business decided that tracking page views and generating additional synthetic page views when a user scrolled would allow the business to measure something that approximates ad impressions. It is all too easy for the business to keep using this subpar measurement methodology because “that’s how it’s always been” when what that company should really do is assess whether it’s possible to track their goal the correct way either merging data in a data warehouse between the frontend data analytics and the backend ad impressions data or just tracking impressions in the web analytics implementation itself.",{"type":47,"tag":55,"props":10540,"children":10542},{"id":10541},"what-about-legacy-processes",[10543],{"type":52,"value":10544},"What About Legacy Processes?",{"type":47,"tag":48,"props":10546,"children":10547},{},[10548],{"type":52,"value":10549},"Just like a legacy KPI, a legacy process is a process that is 2 years or older and due for some scrutiny. Many times, companies keep doing processes because they have become blind to them. If you're over exposed to an advertisement you stop noticing them, if you're exposed to a silly process and told to just keep doing it you eventually forget to even think about the process. At least once per year (or more often if possible) someone should be asking “is there a better way to accomplish this same goal?” about each of your processes. This regular scrutiny helps to clean out the bad processes that otherwise might live on for much too long. Each bad process might only make your company 2% less efficient but if you have enough of them stacking up, they can compound into huge inefficiency.",{"type":47,"tag":55,"props":10551,"children":10553},{"id":10552},"takeaways",[10554],{"type":52,"value":10555},"Takeaways",{"type":47,"tag":522,"props":10557,"children":10558},{},[10559,10564],{"type":47,"tag":141,"props":10560,"children":10561},{},[10562],{"type":52,"value":10563},"Make sure your KPIs focus on your organization's area of growth. Spend less time on old metrics that matter less or may be vanity metrics.",{"type":47,"tag":141,"props":10565,"children":10566},{},[10567],{"type":52,"value":10568},"Make sure you think critically about each of your metrics and processes at least once per year. Ask “why is it done this way or why is it useful?”",{"title":8,"searchDepth":587,"depth":588,"links":10570},[10571,10572,10573,10574],{"id":10502,"depth":587,"text":10505},{"id":10518,"depth":587,"text":10521},{"id":10541,"depth":587,"text":10544},{"id":10552,"depth":587,"text":10555},"content:blog:update-legacy-kpis-and-processes.md","blog/update-legacy-kpis-and-processes.md","blog/update-legacy-kpis-and-processes",1788929328881]