Search

Items tagged with: DataScience


GulliBench is a benchmark that purports to measure the "gullibility" of AI models.

"Current AI models are trained to be extremely good at solving hard problems (i.e., being very smart). But these problems are usually well-defined and have clear solutions. They are created in sterile environments, where data is standardized and things are mostly deterministic. In short, models are trained with the heavy assumption that the data, tools, and knowledge they are handed are undeniably pristine."

"As most of us know, that is not the case in real life. Data is messy and sometimes plain wrong. Tools are buggy and give unreliable results. Assumptions need to be revisited and rewritten."

"GulliBench probes a single, specific failure: taking the data at face value instead of reconciling it against the primary source it should agree with. That's one slice of a much larger category of gullibility failures. Models can be gullible in plenty of other ways: believing a buggy tool's output, accepting a false premise baked into the prompt, deferring to a confident-but-wrong user, following a planted instruction from a document. We don't touch any of that here. We think data-trust is a clean, measurable place to start, but definitely not the whole story."

Here's there most "gullible" top 10:

  1. Opus 5 - 49
  2. Fable 5 - 48
  3. Muse Spark 1.2 - 42
  4. Gemini 3.1 Pro - 23
  5. Kimi K3 - 18
  6. Grok 4.6 - 16
  7. Opus 4.8 - 16
  8. DeepSeek V4 Flash - 16
  9. GLM 5.2 - 15
  10. DeepSeek V4 Pro - 11

GulliBench: Intelligence isn't enough. Measuring skepticism in frontier models.

#solidstatelife #ai #genai #llms #codingai #datascience


Alright Fedi, I'm looking for a job out of AI hell. I know chances aren't great for anyone right now, but I'm taking mine anyway.

I'm a (senior?) #Python developer with 10 years of experience in both #DataScience and #SoftwareEngineering (with a preference for the latter). I can find my way around #Rust and don't shy away from #JavaScript. I've done a bit of #React lately and haven't found it terrible.

I thrive in team environments where I have time both to work on my own and in pairs or mobs ( mob.sh is a great tool I'd use again, but IntelliJ's Code With Me works too).

My greatest passion is automating the tasks around coding away - testing, building, deploying code with either the shortest set of commands or just fully automated from a pipeline on pull/merge requests.

I won't say no to places that require me to use AI, but I have a strong preference to be able to write my code myself.

Would love boosts for reach! Also thanks for reading 💖

#FediJob #BoostsWelcome #BoostsAppreciated #GetFediHired


The media in this post is not displayed to visitors. To view it, please go to the original post.

🎮 Twitch under fire for new gen AI training system that harvests streamer data for Amazon, says it's on by default because 'if it was opt-in, nobody would opt in'

Streams, VODs, comments, and pretty much everything else will go into the gen AI training machine, unless streamers choose to opt out.

📰 Source: Latest from PC Gamer
🔗 Link: pcgamer.com/software/ai/twitch…

#AI #ArtificialIntelligence #DataScience


Your coding agent just spent 10 million tokens to fix a single bug. Is it worth it? 😉

Other researchers have found similar results. Bai et al. (2026) measured coding agents on real software tasks and found they consume roughly 1,000 times the tokens of an ordinary chatbot interaction. And these sort of agents tasks represent the most rapid driver of increased AI usage; Anthropic’s Economic Index found that 97% of their API usage now show “automation-dominant” patterns associated with agents.

#agents #codingagents #greensoftware #sustainability #ai #datascience #research #programming #python #opensource #GreenOps #tech #technology #data #technology #climatechange

theclimatebrink.com/p/the-real…


🏗️ Das neue Interdisziplinäre Zentrum für Modellierung und Simulation (IMoS) ist an die #TUBerlin übergeben worden. Der Forschungsneubau bietet auf 5.700 m² moderne Infrastruktur für interdisziplinäre Spitzenforschung - von #DataScience über Strömungstechnik bis hin zu Geometrie und Visualisierung.

🆕 Zugleich markiert das IMoS den ersten Baustein des neuen Campus Ost der TU Berlin.

Weitere Infos 👉 tu.berlin/go203083/n88870/

#Wissenschaft


The media in this post is not displayed to visitors. To view it, please go to the original post.

From pandemic preparedness to decision-making: #AI is reshaping public health research.

Join the 3rd Artificial Intelligence in Public Health Research Symposium hosted by ZKIPH at #RKI.

📍 Berlin
📆 9–10 Sept 2026

Registration & programme:
🔗 rki.de/ai-symposium-2026

#PublicHealth #AI #DataScience


⁉️ Ihr braucht Umweltdaten für Beruf oder Engagement?

Kommt zum Online-Community Meeting! @fragdenstaat @GreenLegal_EU

Climate HelpDesk und umwelt.info stellen ihre Angebote, auch mit Blick auf die Daten selbst, vor.

💡 Dann seid ihr gefragt: Ideen, Feedback und konkrete Use Cases!

Wann?
17.12.2025 von 10:00 – 12:00 Uhr

Wo?
Online only

Jetzt anmelden:
ki-ideenwerkstatt.de/veranstal…

#opendata #umweltschutz #naturschutz #civiccoding #datascience #science #zivilgesellschaft #bmukn


The media in this post is not displayed to visitors. To view it, please go to the original post.

Today our team member Anna Breger tells her story - “Many little twists and turns have brought me to where I am now and I am absolutely thrilled about my interdisciplinary research project working on image analysis and historical music manuscripts.”

➡️ Find her full story at hermathsstory.eu/anna-breger/

#AppliedMathematics #ImageAnalysis #Music #InterdisciplinaryResearch #NonTraditionalPathways #DataScience #HerMathsStory


The media in this post is not displayed to visitors. To view it, please go to the original post.

⭐ Wir suchen Sie als Wissenschaftl. Mitarbeiter/-in (m/w/d) im Bereich #DataScience

📍 In Frankfurt am Main oder Leipzig

Befristet für die Dauer von 46 Monaten

Bewerben bis zum 14.10.2025

Mehr Infos 👉 bkg.bund.de/SharedDocs/Stellen…

#Job #Karriere #DigitalerZwilling #LiDAR #KI


Any #DataScience people here?

I have a huge #BeaconDB dataset here, recorded with #NeoStumbler.

I would like to play around with it a bit, display densities of radio devices on a map.

I know #QGis, is it reasonably easy to import a CSV file there, assign the coordinates to some columns etc?

Alternatively I know a bit of #R, but #RStudio is #Electron now, so that could get a hassle XD

Btw, there is a #Fedora #COPR for R-Cran packages, how is the situation on #NixOS?


So, eine Englische Vorstellung hab ich, dann fehlt jetzt noch das deutsche #NeuHier.

Während ich auf Englisch Richtung #DataScience und #Python unterwegs bin, geht's im Deutschen eher in den Bereich der #FediEltern und #Verkehrswende. Ich bin viel mit dem #Fahrrad unterwegs - #mdRzA und zur Kita - und möchte mich sicher auf den Straßen bewegen können! Beruflich bin ich sehr interessiert an der #Krankenhausreform, die ist daher auch ab und zu Thema bei mir.